Method and apparatus for detecting defects in a memory having a time-varying bit error rate
By using on-demand defect detection methods with write timestamps and BER in the memory subsystem, the problem of false alarms and low detection efficiency in conventional detection methods is solved, efficient and accurate defect detection is achieved, and system performance is improved.
Patent Information
- Application Number
- CN201980088355.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-10
- Filing Date
- 2019-12-10
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2039-12-10
AI Technical Summary
When existing memory subsystems detect growth defects, conventional defect testing routines generate a large number of false alarms, affecting system performance, and failing to effectively detect defects that grow and manifest at any time during host access.
Using on-demand defect detection method based on write timestamp and bit error rate (BER), we determine whether to trigger the defect detection operation by reading the BER and write timestamp in the read operation, reducing false alarms and improving detection efficiency.
Effectively detect defects in memory components, reduce false alarms, reduce system performance losses, and improve the reliability and stability of memory subsystems.
Smart Images

Figure CN113272905B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to memory subsystems, and more particularly to defect detection in memory components of a memory subsystem having a time-varying bit error rate. Background Art
[0002] A memory subsystem may be a storage system such as a solid state drive (SSD) or a hard disk drive (HDD). The memory subsystem may be a memory module such as a dual in-line memory module (DIMM), a small outline DIMM (SO-DIMM), or a non-volatile dual in-line memory module (NVDIMM). The memory subsystem may include one or more memory components that store data. The memory components may be, for example, non-volatile memory components and volatile memory components. Generally, a host system may utilize the memory subsystem to store data at the memory components and retrieve data from the memory components. Brief Description of the Drawings
[0003] The present disclosure will be more fully understood from the detailed description given below and the accompanying drawings of various embodiments of the present disclosure.
[0004] Figure 1 A exemplary computing environment including a memory subsystem in accordance with some embodiments of the present disclosure is shown.
[0005] Figure 2 is a flowchart of an exemplary method for initiating a defect detection operation to detect a defect in a memory component using a bit error rate (BER) or an error recovery flow (ERF) indicator corresponding to a read operation and a write timestamp in accordance with some embodiments of the present disclosure.
[0006] Figure 3 is a flowchart of an exemplary method for determining whether a W2R delay is within a W2R delay range specified for an initial read voltage level in accordance with some embodiments of the present disclosure.
[0007] Figure 4A is a graph showing the BER varying with the W2R delay for three read voltage levels in accordance with some embodiments of the present disclosure.
[0008] Figure 4B is showing in accordance with some embodiments of the present disclosure Figure 4A a graph of the W2R delay range for a default read level of one of three read voltage levels that is expected to achieve a good BER.
[0009] Figure 5 is a block diagram of a hardware circuit for triggering a defect detection operation in a central processing unit (CPU) of a memory system in accordance with some embodiments of the present disclosure.
[0010] Figure 6 FIG. Figure 6 is a block diagram of an exemplary computer system in which embodiments of the present disclosure may operate. DETAILED DESCRIPTION
[0011] Aspects of the present disclosure relate to defect detection in a memory subsystem having a time-varying bit error rate (BER). The memory subsystem is also referred to herein as a “memory device”. An example of a memory subsystem is a storage device coupled to a central processing unit (CPU) via a peripheral interconnect (e.g., an input / output bus, a storage area network). Examples of storage devices include solid state drives (SSDs), flash drives, universal serial bus (USB) flash drives, and hard disk drives (HDDs). Another example of a memory subsystem is a memory module coupled to the CPU via a memory bus. Examples of memory modules include dual in-line memory modules (DIMMs), small DIMMs (SO-DIMMs), non-volatile dual in-line memory modules (NVDIMMs), etc. The memory subsystem may be, for example, a hybrid memory / storage subsystem. Generally, a host system may utilize a memory subsystem that includes one or more memory components. The host system may provide data to be stored at the memory subsystem and may request retrieval of data from the memory subsystem.
[0012] The memory subsystem may include a plurality of memory components that may store data from the host system. Each memory component may include a different type of media. Examples of media include, but are not limited to, cross-point arrays of non-volatile memory and flash-based memory, such as single-level cell (SLC) memory, triple-level cell (TLC) memory, and quad-level cell (QLC) memory. The characteristics of different types of media may vary from one media type to another. An example of a characteristic associated with a memory component is data density. Data density corresponds to the amount of data (e.g., data bits) that each memory cell of the memory component may store. Taking flash-based memory as an example, a quad-level cell (QLC) may store four bits of data, while a single-level cell (SLC) may store one bit of data. Thus, a memory component that includes QLC memory cells will have a higher data density than a memory component that includes SLC memory cells. Another example of a characteristic of a memory component is access speed. Access speed corresponds to the amount of time for the memory component to access data stored at the memory component.
[0013] Other characteristics of the memory component can be associated with the durability of the memory component in storing data. When data is written to and / or erased from the memory cells of the memory component, the memory cells may be damaged. As the number of write operations and / or erase operations performed on the memory cells increases, the probability that the data stored at the memory cells contains errors increases, and the memory cells are increasingly damaged. The characteristic associated with the durability of the memory component is the number of write operations or the number of program / erase operations performed on the memory cells of the memory component. If the threshold number of write operations performed on the memory cells is exceeded, the data can no longer be reliably stored at the memory cells because the data may contain a large number of errors that cannot be corrected. Different media types can also have different durabilities in storing data. For example, a first media type can have a threshold of 1,000,000 write operations, while a second media type can have a threshold of 2,000,000 write operations. Therefore, the durability of the first media type in storing data is less than that of the second media type in storing data.
[0014] Another characteristic associated with the durability of the memory component in storing data is the total number of bytes written to the memory cells of the memory component. Similar to the number of write operations, as new data is written to the same memory cells of the memory component, the memory cells are damaged, and the probability that the data stored at the memory cells contains errors increases. If the total number of bytes written to the memory cells of the memory component exceeds the threshold of the total number of bytes, the memory cells can no longer reliably store data.
[0015] Another characteristic associated with the memory component is the time-varying BER. In particular, some non-volatile memories (e.g., NAND, phase change, etc.) have a threshold voltage (Vt) distribution that moves over time. Given the same read level, if the Vt distribution moves, the BER changes. Given the Vt distribution at a time instance, there is an optimal read level or an optimal range of read levels that achieves the lowest bit error rate. In particular, the Vt distribution and the BER can vary with the write-to-read (W2R) latency. Due to this time-varying nature of the BER and other noise mechanisms in the memory, a single read level is not sufficient to achieve the optimal memory read BER to meet some system reliability goals. A single read level (e.g., as shown in the three read levels of FIG. 4) achieves a low BER at short W2R latencies, but a high BER at longer latencies. Multiple read levels (e.g., as shown in FIG. 4) can be used in combination to achieve a low BER over the entire W2R latency range.
[0016] Non-volatile memories may have multiple noise mechanisms that increase the BER, such as write wear, interference, defects, etc. However, during error recovery, read retry operations use different read levels to recover data. Read retry operations are used to achieve the lowest BER. For memories with W2R latency-dependent BER, read retry operations are also used to handle a wide range of W2R latencies.
[0017] A particular problem in memory systems is how to detect growing defects. In particular, when NVM-based systems operate over their entire lifetime, defective pages, blocks, and dies may grow. To detect such growing defects (especially those related to read failures), test routines are typically invoked to ensure that high BER or even uncorrectable error correction code (UECC) events are not caused by transient errors. Such test routines can be invoked periodically to detect defects in the system. However, during host access, defects may grow and manifest at any time. This is especially true in very high-performance systems, where multiple accesses to the memory may occur between periodic defect test routines. Additionally, for memories with W2R latency-dependent BER, high BER or read retry events may be mainly workload-induced, meaning that the conventional criteria for triggering defect test routines generate a large number of false alarms, which harms system performance. Conventional memory subsystems typically do not have on-demand trigger criteria for such defect test routines.
[0018] Aspects of the present disclosure address the above and other deficiencies by providing an on-demand trigger criterion for such defect test routines for memories having a time-varying BER, e.g., based on metrics (such as decoder statistics) or based on read retry statistics. In particular, the present disclosure includes innovative methods for defect detection in memories having a time-varying BER, especially a BER that is dependent on the W2R latency. A write timestamp is written to the memory along with the data of each write operation. After each read (possibly with an error recovery process), the system determines whether to trigger a defect test routine based on a combination of its W2R latency and other statistics, including decoder statistics and error recovery process statistics. The present disclosure defines when the test routine can be involved to ensure that high BER or even UECC events are not caused by transient errors. These test routines can be called on demand, rather than periodically as is conventional. Additionally, the present disclosure addresses how to detect defects that grow during host access and manifest at any time. Further, the present disclosure addresses how to reduce false alarms that degrade system performance, as defects can be detected from other events that are primarily workload-induced. That is, the present disclosure minimizes false alarms and reduces performance losses caused by defect management algorithms that run periodically and are triggered by events unrelated to defects. As described herein, the on-demand criterion can be applied to each read operation and can effectively detect abnormally high RBER events to trigger a defect detection algorithm. The trigger criterion can be implemented in hardware, software, or any combination thereof that affects system performance.
[0019] In one embodiment, a processing device performs a read operation to read a data unit that includes data and a write timestamp indicating when the data unit was written to a memory component. The processing device may perform an error recovery process (ERF) to recover the data unit in response to detecting one or more errors in the read operation. The processing device uses the BER corresponding to the read operation and the write timestamp to determine whether to perform a defect detection operation to detect a defect in the memory component. In another embodiment, the processing device uses an indication of the performance of the ERF (also referred to as an ERF indicator) and the write timestamp to determine whether to perform a defect detection operation to detect a defect in the memory component. The performance of the ERF can also be an indication of a defect in the memory component. The processing device initiates a defect detection operation in response to the write timestamp being within a specified range of an initial read voltage level corresponding to the read operation. Additional details of defect detection in a memory component having a time-varying BER are described in more detail below.
[0020] Figure 1FIG. 0 illustrates an exemplary computing environment 100 that includes a memory subsystem 110 in accordance with some embodiments of the present disclosure. The memory subsystem 110 can include media, such as memory components 112A through 112N. The memory components 112A through 112N can be volatile memory components, non-volatile memory components, or a combination of these. In some embodiments, the memory subsystem is a storage system. An example of a storage system is an SSD. In some embodiments, the memory subsystem 110 is a hybrid memory / storage subsystem. Generally, the computing environment 100 can include a host system 120 that uses the memory subsystem 110. For example, the host system 120 can write data to the memory subsystem 110 and read data from the memory subsystem 110.
[0021] The host system 120 can be a computing device such as a desktop computer, laptop computer, network server, mobile device, or such computing devices that include memory and a processing device. The host system 120 can include or be coupled to the memory subsystem 110 such that the host system 120 can read data from or write data to the memory subsystem. The host system 120 can be coupled to the memory subsystem 110 via a physical host interface. As used herein, "coupled to" generally refers to a connection between components, which can be an indirect communication connection or a direct communication connection (e.g., without an intervening component), whether wired or wireless, including connections such as electrical, optical, magnetic, etc. Examples of the physical host interface include but are not limited to Serial Advanced Technology Attachment (SATA) interface, Peripheral Component Interconnect Express (PCIe) interface, Universal Serial Bus (USB) interface, Fibre Channel, Serial Attached SCSI (SAS), etc. The physical host interface can be used to transfer data between the host system 120 and the memory subsystem 110. When the memory subsystem 110 is coupled to the host system 120 via a PCIe interface, the host system 120 can further utilize a Non-Volatile Memory Express (NVMe) interface to access the memory components 112A through 112N. The physical host interface can provide an interface for passing control, address, data, and other signals between the memory subsystem 110 and the host system 120.
[0022] Memory components 112A to 112N can include any combination of different types of non-volatile memory components and / or volatile memory components. An example of a non-volatile memory component includes a NAND-type flash memory. Each of memory components 112A to 112N can include one or more memory cell arrays, such as single-level cells (SLCs) or multi-level cells (MLCs) (e.g., triple-level cells (TLCs) or quad-level cells (QLCs)). In some embodiments, a particular memory component can include an SLC portion and an MLC portion of memory cells. Each of the memory cells can store one or more bits of data (e.g., data blocks) used by host system 120. Although non-volatile memory components such as NAND-type flash memories are described, memory components 112A to 112N can be based on any other type of memory, such as volatile memory. In some embodiments, memory components 112A to 112N can be, but are not limited to, random access memory (RAM), read-only memory (ROM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), phase change memory (PCM), magnetic random access memory (MRAM), negative-or (NOR) flash memory, electrically erasable programmable read-only memory (EEPROM), and cross-point arrays of non-volatile memory cells. A cross-point array of non-volatile memory can perform bit storage together with a stackable cross-grid data access array based on a change in bulk resistance. Additionally, compared with various flash-based memories, cross-point non-volatile memory can perform in-situ write operations, where non-volatile memory cells can be programmed without prior erasure of the non-volatile memory cells. Further, the memory cells of memory components 112A to 112N can be grouped into memory pages or data blocks, which can refer to the units of the memory components for storing data.
[0023] The memory system controller 115 (hereinafter referred to as the "controller") may communicate with the memory components 112A to 112N to perform operations such as reading data, writing data, or erasing data at the memory components 112A to 112N and other such operations. The controller 115 may include hardware, such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The controller 115 may be a microcontroller, application specific logic circuitry (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), or other suitable processor. The controller 115 may include a processor (processing device) 117 configured to execute instructions stored in the local memory 119. In the illustrated example, the local memory 119 of the controller 115 includes embedded memory configured to store instructions for performing various processes, operations, logic flows, and routines for controlling the operation of the memory subsystem 110 (including handling communication between the memory subsystem 110 and the host system 120). In some embodiments, the local memory 119 may include memory registers that store memory pointers, fetched data, etc. The local memory 119 may also include read only memory (ROM) for storing microcode. Although Figure 1 the exemplary memory subsystem 110 in
[0024] has been shown as including the controller 115, in another embodiment of the present disclosure, the memory subsystem 110 does not include the controller 115 and may instead rely on external control (e.g., provided by an external host, or by a processor or controller separate from the memory subsystem).
[0025] The memory subsystem 110 may also include additional circuitry or components not shown. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuitry (e.g., row decoder and column decoder) that may receive addresses from the controller 115 and decode the addresses to access the memory components 112A through 112N.
[0026] The memory subsystem 110 includes a defect detection component 113 that may be used to determine whether to perform a defect detection operation to detect defects in the memory components using BER or ERF indicators and write timestamps in data units, where the write timestamps indicate when the data units were written to the memory components. The defect detection component 113 may trigger a defect detection operation in response to the BER meeting a BER threshold and the calculated W2R (based on the write timestamp) being within the W2R delay range specified for an initial read voltage level. In some embodiments, the controller 115 includes at least a portion of the defect detection component 113. For example, the controller 115 may include a processor 117 (processing device) configured to execute instructions stored in the local memory 119 to perform the operations described herein. In some embodiments, the defect detection component 113 is part of the host system 120, application, or operating system.
[0027] The defect detection component 113 can determine whether the BER corresponding to a read operation meets a threshold criterion when a data unit is read from any one of the memory components 112A to 112N at an initial read voltage level by the read operation. In response to the BER meeting the threshold criterion, the defect detection component 113 can use the BER corresponding to the read operation and the write timestamp to initiate or otherwise perform a defect detection operation to detect defects in the corresponding memory component. For example, after restoring the data unit, the defect detection component 113 can determine whether a reread operation is performed in the ERF. Before performing the ERF, the reread operation is performed at a read voltage level different from the initial read voltage level used with the initial read operation. In response to performing the reread operation in the ERF, the defect detection component 113 can use the BER corresponding to the read operation and the write timestamp to initiate or otherwise perform a defect detection operation to detect defects in the memory component. In another embodiment, the defect detection component 113 can determine whether the ERF has been performed to meet the threshold criterion when a data unit is read from any one of the memory components 112A to 112N at an initial read voltage level by the read operation. In response to the ERF meeting the threshold criterion, the defect detection component 113 can use the indication of the ERF and the write timestamp to initiate or otherwise perform a defect detection operation to detect defects in the corresponding memory component. For example, after restoring the data unit, the defect detection component 113 can determine whether a reread operation is performed in the ERF. Before performing the ERF, the reread operation is performed at a read voltage level different from the initial read voltage level used with the initial read operation. In response to performing the reread operation in the ERF, the defect detection component 113 can use the indication of the ERF and the write timestamp to initiate or otherwise perform a defect detection operation to detect defects in the memory component.
[0028] Figure 2 is a flowchart of an exemplary method 200 for initiating a defect detection operation to detect defects in a memory component using a bit error rate (BER) corresponding to a read operation or an ERF indicator and a write timestamp in accordance with some embodiments of the present disclosure. The method 200 can be performed by processing logic, which can include hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executing on a processing device), or a combination thereof. In some embodiments, the method 200 is performed by Figure 1by the memory defect detection component 113. Although shown in a particular order or sequence, the order of the process can be modified unless otherwise stated. Accordingly, the illustrated embodiments should be understood as merely examples, and the illustrated processes can be performed in different orders, and some processes can be performed in parallel. Additionally, one or more processes can be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other process flows are possible.
[0029] In operation 210, the processing device performs a read operation to read a data unit that includes data and a write timestamp indicating when the data unit was written to the memory component. In operation 220, the processing device detects a high BER condition or an error recovery flow (ERF) condition. The high BER condition can be detected in response to the BER corresponding to the read operation meeting a BER threshold criterion. The ERF condition can be detected when the ERF is performed in response to detecting one or more errors in the read operation to recover the data unit. When the ERF is performed, there can be an indication that the ERF has been performed, such as an ERF indicator. The ERF indicator indicating that the ERF has been performed for the read operation can be used as an indicator of a defect in the memory component. In operation 230, the processing device determines the write-to-read (W2R) delay of the read operation using the current time of the read operation and the write timestamp. In operation 240, the processing device determines whether the W2R delay meets the desired BER condition or ERF condition. In operation 250, the processing device initiates a defect detection operation in response to the W2R delay not meeting the desired BER condition or ERF condition corresponding to the initial read voltage level of the read operation. For example, as shown in FIG. 4, the processing device can store a desired BER range within a specified W2R delay range and can compare the BER and W2R corresponding to the initial read operation with the desired BER range of the specified W2R delay range to determine whether to initiate a defect detection operation. The defect detection operation is initiated in response to the BER being higher than the desired BER range and the W2R delay being within the W2R delay range specified for the initial read voltage level. The defect detection operation is not initiated in response to the BER being within the desired BER range or the W2R delay being outside the W2R delay range specified for the initial read voltage level.
[0030] In another embodiment, after the data unit is recovered by the ERF, the processing device determines whether the BER meets a threshold criterion when the data unit is read at the initial read voltage level by the read operation. In response to the BER meeting the threshold criterion, the processing device uses the BER and the write timestamp to initiate a defect detection operation to detect a defect in the memory component. When the BER does not meet the threshold criterion, the processing device does not initiate a defect detection operation and completes the read operation.
[0031] In another embodiment, after restoring a data unit, the processing device determines whether to perform a reread operation in the ERF. Before performing the ERF, the reread operation is performed at a read voltage level different from the initial read voltage level used with the read operation. In response to performing the reread operation in the ERF, the processing device uses the ERF indicator and the write timestamp to initiate a defect detection operation to detect defects in the memory component. If the reread operation is not performed in the ERF, the processing device does not initiate the defect detection operation and completes the read operation.
[0032] In another embodiment, the processing device determines whether the BER meets a threshold criterion when the data unit is read at the initial read voltage level by the read operation. The processing device determines whether to perform a reread operation in the ERF. As described above, the reread operation is performed at a read voltage level different from the initial read voltage level. In response to the BER meeting the threshold criterion and in response to performing the reread operation in the ERF, the processing device uses the BER and the write timestamp to initiate a defect detection operation to detect defects in the memory component. In response to the BER not meeting the threshold criterion or the reread operation not being performed in the ERF, the processing device does not initiate the defect detection operation and completes the read operation.
[0033] In another embodiment, before performing the ERF, the processing device performs a read operation on a set of memory cells at a first read voltage level to read a data unit in the memory component. As part of the ERF, the processing device performs a reread operation on the set of memory cells at a second read voltage level to restore the data unit. The second read voltage level is different from the first read voltage level. After restoring the data unit, the processing device initiates a defect detection operation to detect defects in the memory component.
[0034] In another embodiment, the processing device uses a default read voltage level to detect one or more errors in a data unit read from a set of memory cells of the memory component. In response to detecting one or more errors in the data unit, as part of the ERF, the processing device performs a reread operation on the set of memory cells at a second read voltage level to restore the data unit. As described above, the second read voltage level is different from the default read voltage level. After restoring the data unit, the processing device initiates a defect detection operation to detect defects in the memory component.
[0035] In another embodiment, the processing device receives a request to write data to the memory component. The processing device obtains a write timestamp and issues a write operation to write the data and the write timestamp as a data unit into the memory component.
[0036] In another embodiment, the processing device obtains a write timestamp and obtains a write temperature value indicating the temperature at the time of writing the data unit. The processing device issues a write operation to write the data, the write timestamp, and the temperature value as a data unit into the memory component. In other embodiments, additional metadata may be stored in association with the write timestamp in the data unit. The metadata may be used in association with defect detection operations.
[0037] In another embodiment, even when ERF is not performed, the processing device may determine to perform a defect detection operation. For example, the original read operation is successful, but the processing device determines that the BER is higher than expected and the W2R is within the W2R delay range specified for the initial read. In such a case, the processing logic may perform a defect detection operation to detect a defect in the memory component.
[0038] In another embodiment, at operation 230, instead of using the BER and the write timestamp, the processing device may use an indication of performing ERF due to the initial read operation being unsuccessful to determine whether to perform a defect detection operation to detect a defect in the memory component, and the W2R (based on the write timestamp in the data unit) is within the W2R delay range specified for the initial read.
[0039] Figure 3 is a flowchart of an exemplary method 300 for determining whether the W2R delay is within the W2R delay range specified for the initial read voltage level according to some embodiments of the present disclosure. Method 300 may be performed by processing logic, which may include hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executing on a processing device), or a combination thereof. In some embodiments, method 300 is performed by Figure 1 the memory defect detection component 113. Although shown in a particular order or sequence, the order of the process may be modified unless otherwise stated. Accordingly, the illustrated embodiments should be understood as merely examples, and the illustrated process may be performed in a different order, and some processes may be performed in parallel. Additionally, one or more processes may be omitted in various embodiments. Thus, not all processes are required in every embodiment. Other process flows are possible.
[0040] At operation 310, the processing device issues a read operation at a specified read voltage level to read a data unit in the memory component. At operation 320, the processing device determines whether the data unit from the read operation is successfully decoded without error. When the processing device determines at operation 320 that the data unit from the read operation is successfully decoded, the processing device determines at operation 325 whether the bit error rate (BER) meets the BER threshold criteria when the data unit is read by the read operation at the specified read voltage level. In response to the BER meeting the BER threshold criteria at operation 325, the processing device determines at operation 350 the write-to-read (W2R) delay between the write timestamp stored in association with the data unit and the original read operation of block 310 using the current time of the initial read operation. At operation 360, the processing device determines whether this W2R delay is expected for this BER condition or the ERF condition. In response to expecting the BER condition or the ERF condition at operation 360, the processing device completes the read operation at operation 360. In response to not expecting the BER condition or the ERF condition at operation 360, the processing device initiates a defect test routine at operation 370.
[0041] For example, the processing device determines at operation 360 whether this given W2R delay is expected to correspond to the BER of the read operation. In response to the given W2R delay not being within the W2R delay range at the initial read voltage level at operation 360, the read operation is completed at operation 330. In response to the given W2R delay being within the W2R delay range at the initial read voltage level at operation 360, the processing device initiates a defect test routine at operation 370. In particular, when the BER of the read operation is higher than the BER range corresponding to the W2R delay range specified for the specified read voltage level (i.e., the acceptable BER range of the W2R delay range serves as the BER threshold criteria) and the given W2R delay is within the W2R delay range at the initial read voltage level, a defect test routine is initiated at operation 370.
[0042] For another example, the processing device determines at operation 360 whether a reread operation is performed in the ERF at operation 340. Before performing the ERF, the reread operation is performed at a read voltage level different from the initial read voltage level used for the read operation of operation 310. In response to a reread operation being performed on the read operation and the given W2R delay being within the W2R delay range at the initial read voltage level, the processing device initiates a defect test routine at operation 370.
[0043] In response to the BER not meeting the BER threshold criteria at operation 325, the processing device completes the read operation at operation 330.
[0044] When the processing device determines in operation 320 that a data unit from a read operation has not been successfully decoded due to an error, the processing device performs an error recovery flow (ERF) in operation 340 to recover the data unit. In some embodiments, during the ERF, the processing device issues one or more reread operations at one or more read voltage levels different from the specified read voltage level. After performing the ERF in operation 340, in operation 350, the processing device determines the W2R latency of the read operation of operation 310 using the current time of the initial read operation and the write timestamp stored in association with the data unit when the data unit was written. It should be noted that the W2R latency is between the write timestamp and the initial read of operation 310 and not any rereads performed during the ERF of operation 340. As described above, in operation 360, the processing device determines whether the W2R latency is within the W2R latency range specified for the specified read voltage level. In response to the W2R latency not being within the W2R latency range in operation 360, the read operation is completed in operation 330. In response to the W2R latency being within the W2R latency range in operation 360, the processing device initiates a defect test routine in operation 370. In particular, when the W2R latency of the read operation is within the W2R latency range specified for the specified read voltage level and an ERF is performed in block 340, a defect test routine is initiated.
[0045] In another embodiment, the processing device obtains a write timestamp and issues a write operation to write the data and the write timestamp as a data unit into the memory component. The processing device can obtain and write a write timestamp for each data unit written into the memory component. The write timestamp can be used to calculate the W2R latency, and the calculated W2R latency can be checked against a corresponding range for the default read voltage level.
[0046] In another embodiment, the processing device also obtains the temperature or other measurements at the time of the write operation and stores the temperature or other measurements as metadata associated with the data. For example, the data unit stores the data, the write timestamp, and the temperature at which the data was written into the memory component.
[0047] In one embodiment, the processing device determines whether a bit error rate (BER) meets a threshold criterion when the data unit is read by a read operation at a specified read voltage level. The processing device initiates a defect test routine in response to the BER meeting the threshold criterion and the W2R latency being within the W2R latency range specified for the specified read voltage level.
[0048] In another embodiment, the processing device determines whether a reread operation is performed in the ERF. The processing device initiates a defect test routine in response to a reread operation being performed in the ERF and the W2R latency being within the W2R latency range specified for the specified read voltage level.
[0049] In another embodiment, the processing device determines whether the BER of an initial read operation meets a threshold criterion and whether to perform a reread operation in the ERF. The processing device initiates a defect test routine in response to meeting both conditions. In particular, the processing device determines whether the BER of the initial read operation meets the threshold criterion when the data unit is read by the read operation at a specified read voltage level. The processing device determines whether to perform a reread operation in the ERF. The processing device initiates a defect test routine in response to the BER meeting the threshold criterion, performing a reread operation in the ERF and the W2R delay being within the W2R delay range specified for the specified read voltage level. In other embodiments, additional checks can be made for other metadata values stored related to the data unit. For example, when a write temperature value is written related to the data unit, the processing device can determine whether a defect test routine should be performed based on considering the W2R delay and the current / write temperature information.
[0050] In another embodiment, the processing device detects one or more errors in a data unit read from a memory component using an initial read voltage level. In response to detecting one or more errors in the data unit, as part of the ERF, the processing device performs a reread operation at a different read voltage level to recover the data unit. After recovering the data unit, the processing device initiates a defect test routine.
[0051] Figure 4A FIG. 400 is a diagram showing the BER varying with the W2R delay for three read voltage levels according to some embodiments of the present disclosure. As described herein, the Vt distribution can shift over time. For example, in the case where the read level (e.g., the second read level (labeled as read level 2) corresponding to the initial read voltage level (also referred to as the default read level)) is the same, if the Vt distribution shifts, then the read voltage level changes over time. Similarly, if the Vt distribution shifts for the first read level, then the bit error rate for this read voltage level changes over time. Similarly, if the Vt distribution shifts for the third read level, then the bit error rate for this read voltage level changes over time. The Vt distribution and the bit error rate can vary with the W2R delay. FIG. 400 shows a bit error rate curve 402 varying with the W2R delay corresponding to the second read level, a bit error rate curve 404 varying with the W2R delay corresponding to the first read level, and a bit error rate curve 406 varying with the W2R delay corresponding to the third read level. Due to the time-varying nature of the BER, a single read level (default read level) is not sufficient to achieve the best memory read BER for system reliability goals. For example, a single read level (e.g., read level 1) achieves a low BER at short W2R delays but a high BER at high delays. Therefore, using multiple read levels (e.g., Figure 4AUsing the embodiments described herein, the W2R delay can be measured using the write timestamp and the current time of the initial read operation to determine whether the measured W2R delay is within the range of about Figure 4B Shown and described are within the specified ranges for specific read levels.
[0052] Figure 4B is a diagram showing some embodiments of the present disclosure Figure 4A 420 of a W2R delay range 408 for a default read level at one of three read voltage levels. If a read is performed at a certain W2R delay within 408, it is expected that the read should achieve a good BER. As described herein, when writing write cells, each data write cell is written to the memory with a write timestamp. Each read operation starts with a default read level. When an uncorrectable error exists, an error recovery process is performed. During the error recovery process, one or more re-read operations are performed at a read level different from the default read level. For example, if Figure 4B As shown in , the default read level is the second read level. The second read level has a bit error rate curve 402 that varies with W2R delay. If it is determined that the decoder statistics (e.g., BER) are high at the default read level, or if a reread operation is triggered with a different read level, the processing device checks the following criteria after recovering the data and the corresponding write timestamp (i.e., successfully decoded by the initial read or using a different read level in the ERF). The check includes measuring the W2R delay of the initial read operation by obtaining the difference between the current time and the write timestamp of the initial read operation and comparing the W2R delay to the W2R delay range 408 specified for the default read level. If the W2R delay of the initial read operation falls within the W2R delay range 408, the defect test routine is triggered; otherwise, the defect test routine is not triggered. It should be noted that the W2R delay is measured for the initial read operation, not for any reread operation that is part of the ERF.
[0053] In one embodiment, the processing device implements this check in a hardware circuit that includes logic circuitry having at least one input, the at least one input being whether the W2R delay is within the W2R delay range 408. The logic circuitry may output an interrupt signal that causes the processing device to perform a defect test routine. In another embodiment, the processing device implements this check in firmware. The firmware calculates the W2R delay and determines whether the W2R delay is within the W2R delay range 408. The firmware may initiate a defect test routine accordingly. In another embodiment, the processing device implements this check as a software routine that is executed in association with a read operation.
[0054] In another embodiment, the processing device may specify a range for each of a plurality of read thresholds. In this way, if a first read level is considered the default read level for an initial read operation, there may be a corresponding W2R delay range for the first read level. Similarly, if a third read level is considered the default read level for an initial read operation, there may be a corresponding W2R delay range for the third read level. It should also be noted that the processing device may include more or fewer than three read levels, and one or more of these multiple read levels may have a W2R delay.
[0055] In another embodiment, a write timestamp may be embedded in the data during a memory write operation, and after each read operation with an ERF, the processing device may determine whether to trigger a defect detection operation based on its W2R delay and other statistics (e.g., the decoding history statistics (BER) of this data unit). In other embodiments, additional metadata may be stored along with the write timestamp. For example, the additional metadata may affect the BER, and the additional metadata may be used in the inspection to determine whether to check for defects based on different combinations of statistics, additional metadata, and the write timestamp.
[0056] The embodiments described herein provide a per-read-operation-on-demand criterion. The embodiments effectively detect anomalous characteristics (e.g., high read bit error rate (RBER) events) and trigger defect detection in response to the write timestamp falling within a specified range designated for the read voltage level used for the initial read operation. The embodiments may minimize false alarms and may reduce performance losses caused by conventional defect management algorithms. The embodiments of the trigger criterion described herein may be simple and may be implemented in hardware without affecting system performance.
[0057] Figure 5It is a block diagram of a hardware circuit 500 that triggers a defect detection operation in a central processing unit (CPU) 510 of a memory system according to some embodiments of the present disclosure. The hardware circuit 500 includes a first comparison circuit system 502, a second comparison circuit system 504, and a logic circuit system 506. The first comparison circuit system 502 may receive the following as inputs: a first signal 512, which indicates a first statistic, such as BER or RBER; and a second signal 514, which indicates a first threshold, such as a BER or RBER threshold. The first comparison circuit system 502 may include one or more comparators to compare the inputs. The first comparison circuit system 502 compares the inputs to generate a first output signal 522, which indicates an abnormal condition, such as a high BER. The second comparison circuit system 504 may receive the following as inputs: a third signal 516, which indicates a first timing statistic, such as a W2R delay; a fourth signal 518, which indicates a lower threshold of a range, such as a W2R delay lower threshold; and a fifth signal 520, which indicates an upper threshold of the range, such as a W2R delay upper threshold. The second comparison circuit system 504 may include one or more comparators to compare the inputs. The second comparison circuit system 504 compares the inputs to generate a second output signal 524, which indicates that the third signal 516 is within the range, such as within the W2R delay range. The logic circuit system 506 may receive the first output signal 522 and the second output signal 524, and output an interrupt 526 to the CPU 510 based on a specific function of the logic circuit system (e.g., an AND function). The interrupt 526 may indicate that a defect detection operation should be performed.
[0058] In one embodiment, the interrupt 526 is the result of the BER of a read operation satisfying the BER threshold criterion and the W2R delay (based on the write timestamp) being within the W2R delay range. The hardware circuit 500 may include different logic and circuit components to determine the conditions for triggering a defect detection operation. For example, the inputs may include a write timestamp and the current time of an initial read operation to calculate the W2R delay before comparing it with the W2R delay range. In other embodiments, the inputs may include other metadata, such as the temperature when the write unit is written to the memory component. Although the logic circuit system 506 is shown as a single AND gate in Figure 5 it, in other embodiments, the logic circuit system 506 may include one or more logic gates that define the function for determining whether a defect detection operation is triggered. Additionally, as described herein, the functions of the hardware circuit 500 may be implemented in firmware or software.
[0059] In another embodiment, a similar comparison and logic circuitry can be used to detect an ERF condition and generate an interrupt to the CPU 510 when the ERF condition is detected. Similarly, other comparison and logic circuitry can be used to detect other conditions that vary with the W2R latency and generate an interrupt to the CPU 510 when the other conditions are detected.
[0060] Figure 6 FIG. shows an example machine of a computer system 600, where a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein. In some embodiments, the computer system 600 can correspond to a host system (e.g., Figure 1 of the memory subsystem 110) that includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 1 of the host system 120) or can be used to operate a controller (e.g., execute an operating system to perform operations corresponding to Figure 1 of the defect detection component 113). In alternative embodiments, the machine can be connected (e.g., networked) to other machines in a LAN, intranet, extranet, and / or the Internet. The machine can operate as a server or client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.
[0061] The machine can be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, network device, server, network router, switch, or bridge or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the machine. Moreover, although a single machine is shown, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein.
[0062] The example computer system 600 includes a processing device 602, a main memory 604 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (e.g., synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), etc.), a static memory 606 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage system 618, which communicate with each other via a bus 630.
[0063] The processing device 602 represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, etc. More particularly, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. The processing device 602 may also be one or more special-purpose processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. The processing device 602 is configured to execute the instructions 626 to perform the operations and steps discussed herein. The computer system 600 may further include a network interface device 608 to communicate via the network 620.
[0064] The data storage system 618 may include a machine-readable storage medium 624 (also referred to as a computer-readable storage medium) having stored thereon one or more sets of instructions 626 or software embodying any one or more of the methods or functions described herein. The instructions 626 may also reside, completely or at least partially, within the main memory 604 and / or the processing device 602 during execution by the computer system 600, the main memory 604 and the processing device 602 also constituting machine-readable storage media. The machine-readable storage medium 624, the data storage system 618, and / or the main memory 604 may correspond to Figure 1 the memory subsystem 110.
[0065] In one embodiment, the instructions 626 include instructions for implementing the functionality corresponding to an ERF component (e.g., Figure 1 the defect detection component 113). Although the machine-readable storage medium 624 is shown as a single medium in the exemplary embodiment, the term "machine-readable storage medium" or "computer-readable storage medium" should be considered to include a single medium or multiple media storing one or more sets of instructions. The term "machine-readable storage medium" should also be considered to include any medium that is capable of storing or encoding a set of instructions executable by a machine and that causes the machine to perform any one or more of the methods of the present disclosure. Thus, the term "machine-readable storage medium" should be considered to include, but not be limited to, solid-state memory, optical media, and magnetic media.
[0066] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of operations leading to a desired result. The operations are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. Sometimes, for the sake of generality, it has proven convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc.
[0067] However, it should be borne in mind that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may refer to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical (electronic) quantities within the registers and memories of the computer system and transforms them into other data similarly represented as physical quantities within the computer system memory or registers or other such information storage systems.
[0068] The present disclosure also relates to an apparatus for performing the operations herein. The apparatus may be specially constructed for the intended purposes, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer readable storage medium, such as but not limited to any type of disk, including floppy disks, optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of media suitable for storing electronic instructions, each coupled to the computer system bus.
[0069] The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. Various general purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized apparatus to perform the method. The structure of various of these systems will be set forth as will be described below. In addition, the present disclosure has been described without reference to any particular programming language. It will be understood that a variety of programming languages may be used to implement the teachings of the present disclosure as described herein.
[0070] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, and the instructions may be used to program a computer system (or other electronic device) to perform the processes according to the present disclosure. The machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, the machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory (“ROM”), a random access memory (“RAM”), a magnetic disk storage medium, an optical storage medium, a flash memory component, etc.
[0071] In the foregoing specification, embodiments of the present disclosure have been described with reference to specific exemplary embodiments thereof. Obviously, various modifications can be made thereto without departing from the broader spirit and scope of the embodiments of the present disclosure set forth in the following claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
Claims
1. A system, comprising: a memory component; and a processing device operatively coupled to the memory component to: perform a read operation at a current time to read a data unit, the data unit including data and a write timestamp indicating when the data unit was written to the memory component; determine whether an error recovery flow (ERF) condition is detected, wherein the ERF condition is detected in response to the ERF performing to recover the data unit in response to detecting one or more errors in the read operation; determine whether a bit error rate (BER) condition is detected, wherein the BER condition is detected in response to the BER corresponding to the read operation meeting a threshold criterion; determine a write-to-read (W2R) delay of the read operation, wherein the W2R delay includes a difference between the current time of the read operation and the write timestamp indicating when the data unit was written to the memory component; in response to detecting at least one of the ERF condition or the BER condition, determine whether the W2R delay is within a W2R delay range corresponding to an initial read voltage level used by the read operation to read the data unit; and initiate a defect detection operation in response to the W2R delay being within the W2R delay range, the defect detection operation being for detecting a time-varying defect in the memory component.
2. The system according to claim 1, wherein after the data unit is recovered by the ERF, the processing device is further configured to: determine whether the BER meets the threshold criterion when the data unit is read by the read operation at the initial read voltage level; and initiate the defect detection operation to detect the defect in the memory component in response to the BER meeting the threshold criterion.
3. The system according to claim 1, wherein after the data unit is recovered by the ERF, the processing device is further configured to: determine whether a reread operation is performed in the ERF, wherein the reread operation is performed at a read voltage level different from the initial read voltage level used in the read operation; and initiate the defect detection operation to detect the defect in the memory component in response to the reread operation being performed in the ERF.
4. The system according to claim 1, wherein the processing device is further configured to: determine whether the BER meets the threshold criterion when the data unit is read by the read operation at the initial read voltage level; determine whether a reread operation is performed in the ERF, wherein the reread operation is performed at a read voltage level different from the initial read voltage level; and initiate the defect detection operation to detect the defect in the memory component in response to the BER meeting the threshold criterion and in response to the reread operation being performed in the ERF.
5. The system according to claim 1, wherein the processing device is further configured to: Before performing the ERF, perform the read operation on a plurality of memory cells at the initial read voltage level to read the data unit in the memory component; As part of the ERF, perform a reread operation on the plurality of memory cells at a second read voltage level to recover the data unit, wherein the second read voltage level is different from the initial read voltage level; and After recovering the data unit, initiate the defect detection operation to detect the defect in the memory component.
6. The system according to claim 1, wherein the processing device is further configured to: Detect one or more errors in the data units read from the plurality of memory cells of the memory component using the initial read voltage level; and In response to detecting one or more errors in the data unit, as part of the ERF, perform a reread operation on the plurality of memory cells at a second read voltage level to recover the data unit, wherein the second read voltage level is different from the initial read voltage level; and After recovering the data unit, initiate the defect detection operation to detect the defect in the memory component.
7. The system according to claim 1, wherein the processing device is further configured to: Obtain the write timestamp; and Issue a write operation to write the data and the write timestamp as the data unit into the memory component.
8. The system according to claim 1, wherein the processing device is further configured to: Obtain the write timestamp; Obtain a write temperature value indicating the temperature at which the data unit was written; and Issue a write operation to write the data, the write timestamp, and the temperature value as the data unit into the memory component.
9. A method, comprising: Issue a read operation at a specified read voltage level at the current time to read a data unit in a memory component; Determine that the data unit from the read operation was not successfully decoded due to an error; Perform an error recovery flow (ERF) to recover the data unit, wherein performing the ERF includes issuing a reread operation at a read voltage level different from the specified read voltage level; Determine the write-to-read (W2R) latency of the read operation, wherein the W2R latency includes the difference between the current time of the read operation and a write timestamp indicating when the data unit was written into the memory component; Determine whether the W2R latency is within a W2R latency range corresponding to the specified read voltage level; and In response to the W2R latency being within the W2R latency range, initiate a defect test routine for testing time-varying defects in the memory component.
10. The method according to claim 9, further comprising: Obtain the write timestamp; and Issue a write operation to write the data and the write timestamp as the data unit into the memory component.
11. The method according to claim 9, further comprising: Determine whether a bit error rate (BER) corresponding to the read operation meets a threshold criterion when a data unit is read by the read operation at the specified read voltage level, wherein initiating the defect test routine includes initiating the defect test routine in response to the BER meeting the threshold criterion.
12. The method according to claim 9, further comprising: Determine whether the reread operation is performed in the ERF, and wherein initiating the defect test routine includes initiating the defect test routine in response to the reread operation being performed in the ERF and the W2R delay being within the W2R delay range specified for the specified read voltage level.
13. The method according to claim 9, further comprising: Determine whether a bit error rate (BER) corresponding to the read operation meets a threshold criterion when a data unit is read by the read operation at the specified read voltage level; and Determine whether the reread operation is performed in the ERF, and wherein initiating the defect test routine includes initiating the defect test routine in response to the BER meeting the threshold criterion, the reread operation being performed in the ERF, and the W2R delay being within the W2R delay range specified for the specified read voltage level.
14. The method according to claim 9, further comprising: Detect one or more errors in the data unit read from the memory component using the specified read voltage level; and In response to detecting one or more errors in the data unit, perform the reread operation at a different read voltage level as part of the ERF to recover the data unit, wherein initiating the defect test routine includes initiating the defect test routine after recovering the data unit.
15. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to: Issue a read operation at a specified read voltage level at a current time to read a data unit in a memory component; Determine that the data unit from the read operation was not successfully decoded due to an error; Perform an error recovery flow (ERF) to recover the data unit, wherein performing the ERF includes issuing a reread operation at a read voltage level different from the specified read voltage level; Determine a write-to-read (W2R) delay of the read operation, wherein the W2R delay includes a difference between the current time of the read operation and a write timestamp indicating when the data unit was written to the memory component; Determine whether the W2R delay is within a W2R delay range corresponding to the specified read voltage level; and In response to the W2R delay being within the W2R delay range, initiate a defect test routine for testing time-varying defects in the memory component.
16. The non-transitory computer-readable storage medium according to claim 15, wherein the processing device is further configured to: Obtain the write timestamp; and Issue a write operation to write the data and the write timestamp as the data unit into the memory component.
17. The non-transitory computer-readable storage medium according to claim 15, wherein the processing device is further configured to: Determine whether a bit error rate (BER) corresponding to the read operation meets a threshold criterion when a data unit is read by the read operation at the specified read voltage level, wherein the defect test routine is initiated in response to the BER meeting the threshold criterion.
18. The non-transitory computer-readable storage medium according to claim 15, wherein the processing device is further configured to: Determine whether the reread operation is performed in the ERF, and wherein the defect test routine is initiated in response to the reread operation being performed in the ERF and the W2R delay being within the W2R delay range specified for the specified read voltage level.
19. The non-transitory computer-readable storage medium according to claim 15, wherein the processing device is further configured to: Determine whether a bit error rate (BER) corresponding to the read operation meets a threshold criterion when a data unit is read by the read operation at the specified read voltage level; and Determine whether the reread operation is performed in the ERF, and wherein the defect test routine is initiated in response to the BER meeting the threshold criterion, the reread operation being performed in the ERF, and the W2R delay being within the W2R delay range specified for the specified read voltage level.
20. The non-transitory computer-readable storage medium according to claim 15, wherein the processing device is further configured to: Detect one or more errors in the data unit read from the memory component using the specified read voltage level; and In response to detecting one or more errors in the data unit, perform the reread operation at the different read voltage level as part of the ERF to recover the data unit, wherein the defect test routine is initiated after recovering the data unit.
Citation Information
Patent Citations
Mitigating read errors following programming in a multi-level non-volatile memory
US10101931B1
Reference voltage calibration in memory during runtime
US20170358369A1