Method and system for managing block retirement for temporary operating conditions
By detecting data loss in memory component blocks, identifying behavioral criteria, and performing stress tests, the problem of improper block exit management under temporary operating conditions in the prior art is solved, and higher memory subsystem reliability and performance are achieved.
Patent Information
- Application Number
- CN202011132140.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-10-22
- Filing Date
- 2020-10-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-10-21
AI Technical Summary
The existing memory subsystem is difficult to effectively manage block withdrawal under temporary operating conditions, which may lead to the healthy block being marked as withdrawal prematurely, affecting the performance and reliability of the memory subsystem.
By detecting data loss in memory component blocks, identifying associated codes of behavior, incrementing the counter associated with the block, responding that the counter value reaches a threshold, specifying the block as an isolated block, performing stress tests, and deciding whether to exit the block based on the test results.
This method can effectively reduce block withdrawal caused by temporary operating conditions, protect healthy blocks, improve the reliability and performance of the memory subsystem, and prevent unnecessary block withdrawals.
Smart Images

Figure CN112700815B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to memory subsystems, and more particularly, to managing block retirements for temporary operating conditions. Background Art
[0002] The memory subsystem may include one or more memory components that store data. The memory components may be, for example, non-volatile memory components and volatile memory components. In general, the host system may utilize the memory subsystem to store data at the memory components and retrieve data from the memory components. Summary of the invention
[0003] According to aspects of the present application, a system is provided. The system includes: a memory component; and a processing device operatively coupled to the memory component to perform the following operations: detecting a data loss occurrence in a block of the memory component; identifying a behavior criterion associated with the data loss occurrence in the block of the memory component; incrementing a counter associated with the block in response to the occurrence of the behavior criterion, wherein the value of the counter corresponds to the number of occurrences of a plurality of behavior criteria associated with the data loss occurrence in the block; and in response to determining that the value of the counter satisfies a first threshold criterion: designating the block as an isolated block, performing a stress test of a plurality of stress tests on the block, and in response to the block failing a first stress test, retiring the block of the memory component.
[0004] According to another aspect of the present application, a method is provided. The method includes: detecting, by a processing device, a data loss occurrence in a block of a memory component; identifying a behavior criterion associated with the data loss occurrence in the block of the memory component; incrementing a counter associated with the block in response to the occurrence of the behavior criterion, wherein the value of the counter corresponds to the number of occurrences of a plurality of behavior criteria associated with the data loss occurrence in the block; and in response to determining that the value of the counter satisfies a first threshold criterion: designating the block as an isolated block, performing a stress test of a plurality of stress tests on the block, and in response to the block failing a first stress test, retiring the block of the memory component.
[0005] According to another aspect of the present application, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium includes instructions that, when executed by a processing device, cause the processing device to perform the following operations: identify an operating condition associated with the occurrence of data loss in a block of a memory component; and in response to the operating condition satisfying a behavior criterion: designate the block as an isolated block, perform a stress test of a plurality of stress tests on the block, and in response to the block failing a first stress test, retire the block of the memory component. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The present disclosure will be more fully understood from the detailed description given below and the accompanying drawings of various embodiments of the present disclosure.
[0007] Figure 1 An example computing environment including a memory subsystem according to some embodiments of the present disclosure is described.
[0008] Figure 2 is a schematic sequence diagram illustrating various states that a block may occupy at a given point in time according to some embodiments of the present disclosure.
[0009] Figure 3 is a flow chart of an example method for managing block retirement for temporary operating conditions according to some embodiments of the present disclosure.
[0010] Figure 4 is a flow chart of an example method for adding a block that causes data loss to a watch list before retiring the block, according to some embodiments of the present disclosure.
[0011] Figure 5 is a flow chart of an example method for adding a block causing data loss to a quarantine queue before retiring the block, according to some embodiments of the present disclosure.
[0012] Figure 6 is a block diagram of an example computer system in which embodiments of the present disclosure may operate. DETAILED DESCRIPTION
[0013] Aspects of the present disclosure are directed to systems and methods for managing block retirement for temporary operating conditions in a memory subsystem. The memory subsystem may be a storage device, a memory module, or a mixture of a storage device and a memory module. Figure 1 Examples of storage devices and memory modules are described. In general, a host system may utilize a memory subsystem that includes one or more memory components (e.g., a memory device that stores data). The host system may provide data for storage at the memory subsystem and may request retrieval of data from the memory subsystem.
[0014] The memory device may be a non-volatile memory device. A non-volatile memory device is a package of one or more dies. Each die may be composed of one or more planes. The planes may be grouped into logical units (LUNs). For some types of non-volatile memory devices (e.g., NAND devices), each plane is composed of a set of physical blocks. Each block is composed of a set of pages. Each page is composed of a set of memory cells storing data bits. The memory subsystem is expected to operate within certain specifications designed to allow proper storage and retrieval of host data at the memory subsystem. When the memory subsystem operates outside the design specifications, certain blocks may return errors in the form of data loss. When a read operation of a previously written data bit to a block fails and a subsequent system-level error handling flow fails to recover the data, data loss of a block may occur. In one example, when a customer (e.g., when testing the memory subsystem) violates the design specifications of the memory subsystem, the memory subsystem may experience data loss. For example, the memory subsystem may operate at an operating temperature that violates the design specifications, which may cause a cross-temperature failure to occur. Cross-temperature failures can occur when the memory subsystem is operated in an environment with widely varying temperatures such that memory cells are programmed at a first temperature and later read at a significantly different temperature. In another example, a design specification violation can occur when the memory subsystem is stored at an abnormal power-off storage temperature for an extended period of time (e.g., a memory drive is stored at 55 degrees Celsius for several months when the allowable power-off storage temperature is 30 degrees Celsius).
[0015] Because data loss caused by design specification violations cannot be recovered by the error handling mechanism, the memory subsystem may retire blocks that suffer data loss because it is presumed that the data loss is an indication of a defective block. A "retired block" refers to a block that is permanently marked as unusable, making it unavailable to the memory subsystem for storing host data during the life of the block. Therefore, when the memory component itself is otherwise healthy and non-defective, retiring a block due to temporary operating conditions, as is the case when the memory subsystem operates under conditions that violate the design specification, may not be desirable. Therefore, the block retirement mechanism needs to incorporate a mitigation process so that data loss due to temporary abnormal operating conditions does not cause healthy blocks to retire.
[0016] Conventionally, the common practice of managing block retirement includes executing a system-level error handling program to try to recover data loss before retiring the block. For example, if the error handling program is able to recover data loss in the block, the block may not be marked as needing to retire. If, on the other hand, the error handling program is unable to recover data loss, the block may be marked as needing to retire. Although in some cases, the use of an error handling program can slow down block retirement, the error handling program may still be unable to recover data loss caused by temporary operating conditions, and therefore may not be able to slow down block retirement due to temporary operating conditions even when the block itself is not inherently defective. For example, when the memory subsystem operates at a crossover temperature that violates the design specification, the block of the memory subsystem may experience data loss that triggers the error handling program. When trying to recover data loss, the error handling program may not be able to recognize that the current abnormal crossover temperature is temporary, and therefore, if the error handling program fails to recover data loss, the block may be marked as retired. Retiring the block in this case may not be desirable, because the block can be correctly executed when operating at a normal crossover temperature that complies with the design specification. Therefore, different techniques for managing block retirement may be preferred to improve performance and reduce data loss while ensuring that healthy blocks are not retired prematurely.
[0017] Aspects of the present disclosure address the above and other deficiencies by implementing systems and methods for managing block retirement for temporary operating conditions. The memory subsystem may initially detect the occurrence of data loss in the block that cannot be recovered by the system-level error handling program. The memory subsystem may then identify the behavior criteria that caused the data loss in the block to occur. In certain embodiments, the behavior criteria may be whether one of several failure mechanisms meets a predetermined threshold criterion. The failure mechanism may include abnormal operating conditions related to the crossover temperature of the block, program / erase cycles (hereinafter P / E cycles), read interference, or a combination thereof. Managing block retirement under temporary operating conditions can ensure that the drive remains functional when design specifications are violated. It can further protect against write protection effects caused by accidental use anomalies attributed to customers. A write protection state occurs due to uncorrectable errors when the memory subsystem operates in a read-only condition. In the case of failure to properly manage block retirement events, the probability of the memory subsystem entering a write protection state may increase the write protection state. In addition, managing block retirement under temporary operating conditions can also prevent minor firmware vulnerabilities (bugs) that can trigger unnecessary block retirement events.
[0018] In certain embodiments, once a behavior criterion has been identified, the memory subsystem may add the block to a watch list so that the block may be monitored for additional data loss events. In one example, in response to the occurrence of a behavior criterion, the block may be added to a watch list by incrementing a watch list counter associated with the block. While on the watch list, the memory block may be monitored for abnormal behavior (e.g., additional data loss events) while continuing to be available to the memory subsystem for storing host data. For example, each time a data loss event occurs due to an abnormal crossover temperature of a memory cell corresponding to the block, the watch list counter for the block may be incremented. In response to determining that the value of the watch list counter for the block exceeds a predetermined threshold, the memory subsystem may designate the block as a quarantine block by adding the block to a quarantine queue.
[0019] The memory subsystem may maintain an isolation queue to perform certain stress tests on the isolated blocks to evaluate the health of the isolated blocks. The stress test of the isolated blocks may include a cross temperature test, a P / E cycle test, a read disturbance test, or a combination thereof of the blocks. When in the isolation queue, the block may be designated as unavailable for use by the memory component to store host data within a predetermined time period. After the block is designated as isolated, the memory subsystem may perform a stress test of the block using test data to evaluate the health of the block based on the result of the stress test. As will be described in more detail herein, in one embodiment, if the stress test result of the block does not meet the test criteria, the memory subsystem may determine that the block is defective and may retire the block. If, on the other hand, the memory subsystem determines that the first stress test result meets the test criteria, the memory subsystem may designate the block as a non-isolated block, thereby allowing the block to exit the isolation queue. The block may then be considered a healthy block by the memory subsystem and may be used to store host data.
[0020] Figure 1An example computing environment 100 including a memory subsystem 110 according to some embodiments of the present disclosure is illustrated. The memory subsystem 110 may include media, such as memory components 112A-112N (hereinafter also referred to as "memory devices"). The memory components 112A-112N may be volatile memory components, non-volatile memory components, or a combination of such components. The memory subsystem 110 may be a storage device, a memory module, or a hybrid of a storage device and a memory module. Examples of storage devices include solid-state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controller (eMMC) drives, universal flash storage devices (UFS) drives, and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small outline DIMMs (SO-DIMMs), and non-volatile dual in-line memory modules (NVDIMMs).
[0021] The computing environment 100 may include a host system 120 coupled to a memory system. The memory system may include one or more memory subsystems 110. In some embodiments, the host system 120 is coupled to memory subsystems 110 of different types. Figure 1 An example of a host system 120 coupled to one memory subsystem 110 is illustrated. The host system 120 uses, for example, the memory subsystem 110 to write data to the memory subsystem 110 and read data from the memory subsystem 110. As used herein, "coupled to" generally refers to a connection between components, which may be an indirect communication connection or a direct communication connection (e.g., without intervening components), whether wired or wireless, including connections such as electrical connections, optical connections, magnetic connections, etc.
[0022] The host system 120 may be a computing device such as a desktop computer, a portable computer, a network server, a mobile device, an embedded computer (e.g., a computer included in a vehicle, an industrial device, or a networked commercial device), or such computing devices that include a memory and a processing device. The host system 120 may include or be coupled to the memory subsystem 110 so that the host system 120 can read data from the memory subsystem 110 or write data to the memory subsystem 110. The host system 120 may be coupled to the memory subsystem 110 via a physical host interface. Examples of the physical host interface include, but are not limited to, a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Universal Serial Bus (USB) interface, a Fibre Channel, a Serial Attached SCSI (SAS), etc. The physical host interface may be used to transmit data between the host system 120 and the memory subsystem 110. When the memory subsystem 110 is coupled to the host system 120 through a PCIe interface, the host system 120 may also utilize an NVM Express (NVMe) interface to access the memory components 112A to 112N. The physical host interface may provide an interface for transferring control, address, data, and other signals between the memory subsystem 110 and the host system 120 .
[0023] Memory components 112A to 112N may include any combination of different types of non-volatile memory components and / or volatile memory components. Examples of non-volatile memory components include NAND type flash memory. Each of memory components 112A to 112N may include one or more arrays of memory cells, such as single-level cells (SLC), multi-level cells (MLC), triple-level cells (TLC), or quad-level cells (QLC). In some embodiments, a particular memory component may include both an SLC portion and an MLC portion of a memory cell. Each of the memory cells may store one or more data bits for use by the host system 120. Although non-volatile memory components such as NAND type flash memory are described, memory components 112A to 112N may be based on any other type of memory, such as volatile memory. In some embodiments, the memory components 112A to 112N may be, but are not limited to, random access memory (RAM), read-only memory (ROM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), phase change memory (PCM), magnetic random access memory (MRAM), or non-(NOR) flash memory, electrically erasable programmable read-only memory (EEPROM), and a cross-point array of non-volatile memory cells. The cross-point array of non-volatile memory may be combined with a stackable cross-grid data access array to perform bit storage based on changes in body resistance. In addition, compared to many flash-based memories, cross-point non-volatile memories may perform write-in-place operations, where non-volatile memory cells may be programmed without pre-erasing the non-volatile memory cells. Furthermore, the memory cells of the memory components 112A to 112N may be grouped into memory pages or blocks, which may refer to cells of a memory component used to store data. The data blocks may be further grouped into one or more planes on each of the memory components 112A-112N, where operations may be performed on each of the planes simultaneously. Corresponding blocks from different planes may be associated with each other in the form of stripes, spanning multiple planes.
[0024] The memory system controller 115 (hereinafter referred to as the "controller") can communicate with the memory components 112A to 112N to perform operations, such as reading data, writing data, or erasing data at the memory components 112A to 112N, and other such operations. The controller 115 may include hardware, such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The controller 115 may be a microcontroller, a dedicated logic circuit (e.g., a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc.), or other suitable processor. The controller 115 may include a processor (processing device) 117 configured to execute instructions stored in a local memory 119. In the illustrated example, the local memory 119 of the controller 115 includes an embedded memory configured to store instructions for executing various processes, operations, logic flows, and routines that control the operation of the memory subsystem 110, including handling communications between the memory subsystem 110 and the host system 120. In some examples, the local memory 119 may include memory registers that store memory pointers, extracted data, etc. Local memory 119 may also include read-only memory (ROM) for storing microcode. Figure 1 The example memory subsystem 110 in FIG. 1 has been illustrated as including a controller 115, but in another embodiment of the present disclosure, the memory subsystem 110 may not include a controller 115 and may instead rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem).
[0025] In general, the controller 115 may receive commands or operations from the host system 120 and may convert the commands or operations into instructions or suitable commands to achieve the desired access to the memory components 112A to 112N. The controller 115 may be responsible for other operations such as wear leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address translation between logical block addresses and physical block addresses associated with the memory components 112A to 112N. The controller 115 may additionally include host interface circuitry to communicate with the host system 120 via a physical host interface. The host interface circuitry may convert commands received from the host system into command instructions to access the memory components 112A to 112N, and convert responses associated with the memory components 112A to 112N into information for the host system 120.
[0026] The memory subsystem 110 may also include additional circuitry or components not shown. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuits (e.g., row decoders and column decoders) that may receive addresses from the controller 115 and decode the addresses to access the memory components 112A to 112N.
[0027] In one embodiment, the memory subsystem 110 includes a block retirement component 113, which can be used to manage the retirement of blocks in one or more of the memory components 112A to 112N of the memory subsystem 110 by monitoring and testing the operating conditions of the blocks. In one embodiment, the block retirement component 113 can detect the occurrence of data loss in the block that cannot be recovered by the system-level error handling program. The block retirement component 113 can then identify the behavior criteria that caused the data loss in the block to occur. In certain embodiments, the behavior criteria can be whether one of several failure mechanisms meets a predetermined threshold criterion. For example, if the failure mechanism exceeds a predetermined threshold, then it meets the threshold criterion. Similarly, when the failure mechanism meets the predetermined threshold criterion, the failure mechanism exceeds the predetermined threshold. The failure mechanism may include abnormal operating conditions related to the crossover temperature of the block, P / E cycles, read interference, or a combination thereof. When the memory subsystem is operated in an environment with widely varying temperatures so that the memory cells are programmed at a given temperature and later read at a significantly different temperature, a crossover temperature failure may occur. Problems may arise over several P / E cycles when memory cells within a memory subsystem lose charge over time due to multiple P / E cycles operated on the respective memory cells. Read disturb problems occur when a read of a particular location of a memory subsystem (e.g., one row of memory cells of a block) affects the threshold voltage of an unread adjacent location (e.g., a different row of the same block).
[0028] Upon identifying the behavior criteria, the block retirement component 113 may add the block to a watch list so that the block can be monitored for additional data loss events. In one example, the block retirement component 113 may add the block to a watch list by incrementing a watch list counter associated with the block in response to the occurrence of the behavior criteria. In response to determining that the value of the watch list counter for the block satisfies a predetermined threshold criterion, the block retirement component 113 may designate the block as a quarantined block by adding the block to a quarantine queue and marking the block as unavailable for use by the memory subsystem 110 to store host data.
[0029] The block retirement component 113 may maintain an isolation queue to perform certain stress tests on the isolated blocks to evaluate the health of the isolated blocks. The stress test of the isolated blocks may include a cross temperature test of the block, a P / E cycle test, a read disturbance test, or a combination thereof. After the block is designated as isolated, the block retirement component 113 may perform a stress test of the block to evaluate the health of the block based on the result of the stress test. In one embodiment, the test criteria may compare the first stress test result of the block with the second stress test result of the healthy block and determine whether the first stress test result is within the expected variance with the second stress test result of the healthy block. If the first stress test result of the block does not meet the test criteria, the block retirement component 113 may retire the block. If on the other hand, the block retirement component 113 determines that the first stress test result meets the test criteria, the block retirement component 113 may designate the block as a non-isolated block, thereby allowing the block to exit the isolation queue. The block may then be considered a healthy block by the memory subsystem 110 and may be used to store host data.
[0030] Figure 2 2 is a schematic sequence diagram 200 illustrating various states that a block may occupy at a given time according to some embodiments of the present disclosure. In one embodiment, Figure 1 Each memory component 112A-N may contain hundreds of blocks. A block 210 may be assigned a health state 220 when first opened by the memory subsystem 110. In certain embodiments, the health state 220 may indicate that the block is not marked by the memory subsystem 110 as being monitored in a watch list or being stress tested in an isolation queue. At operation 224, the block 210 may experience a data loss event that may be detected by the memory subsystem 110. Data loss may occur when a read operation of a previously written data bit to a block fails and a subsequent system-level error handling flow fails to recover the data. In one example, the memory subsystem 110 may experience data loss when a customer (e.g., when testing the memory subsystem) violates a design specification of the memory subsystem. In one embodiment, the memory subsystem may operate at a crossover temperature that violates the design specification, which causes a crossover temperature failure to occur. After detecting data loss in the block 210, the memory subsystem 110 may identify a behavior criterion that caused the data loss in the block 210 to occur. In certain embodiments, the behavior criterion may be whether the failure mechanism satisfies a predetermined threshold criterion. Failure mechanisms may include abnormal operating conditions related to crossover temperature of block 210, P / E cycling, read disturb, or a combination thereof.
[0031] The memory subsystem 110 may then assign the block 210 to the "on watch list" state 230. The "on watch list" state 230 causes the memory subsystem 110 to monitor the block 210 for further occurrences of behavior criteria that cause data loss events, or other behavior criteria. In one example, the memory subsystem 110 may assign the block 210 to the "on watch list" state 230 by incrementing a watch list counter associated with the block 210 at operation 233. The watch list counter associated with the block 210 may be incremented for each occurrence of the behavior criteria of the block 210. Furthermore, the block 210 may continue to be available to the memory subsystem 110 for storing host data while in the "on watch list" state 230. For example, the block 210 may be assigned the "on watch list" state 230 due to abnormal crossover temperature of a memory cell corresponding to the block 210. During the next power-on event of the memory subsystem 110, the crossover temperature of the memory cells may be measured and if an abnormal crossover temperature is again detected, the watch list counter of block 210 may be incremented again.
[0032] At operation "appears in watch list more often than other blocks" 252, the memory subsystem 110 may detect that block 210 appears in the "in watch list" state 230 at a significantly higher frequency than other blocks within the memory subsystem 110 (e.g., by comparing the watch list counter associated with block 210 with the watch list counters of other blocks within the memory subsystem 110). The memory subsystem 110 may then determine that the data loss in block 210 is not due to an abnormal operating condition of the memory subsystem 110, since other blocks within the subsystem 110 do not experience the same rate of data loss. The memory subsystem 110 may then determine that block 210 is unhealthy and may mark block 210 as retired by assigning it a retired state 250. When block 210 is in the retired state 250, it may be presumed to be a defective block 260 for the duration of its life, and the memory subsystem 110 may not use block 210 for storing host data.
[0033] At operation “Counter satisfies threshold criteria” 235, the memory subsystem 110 may detect that the block 210 appears in the “in watch list” state 230 at a number of times that the predetermined threshold criteria are satisfied (e.g., by consulting a watch list counter associated with the block 210). For example, when the watch list counter exceeds a predetermined threshold, the value of the watch list counter may satisfy the threshold criteria. In response to determining that the value of the watch list counter of the block 210 satisfies the predetermined threshold criteria, the memory subsystem 110 may assign the block 210 an isolation state 240. In one example, the memory subsystem 110 may assign the block 210 an isolation state 240 by adding the block 210 to an isolation queue. The memory subsystem 110 may maintain an isolation queue to perform certain stress tests on the isolated blocks to assess the health of the isolated blocks. The stress tests of the isolated blocks may include cross temperature tests, P / E cycle tests, read disturbance tests, or combinations thereof of the blocks. In one example, the stress tests of the isolated blocks may be run during idle time of the memory subsystem 110 to avoid introducing latency to the memory drive. Additionally, while in the quarantine state 240 , the block 210 may be designated as unavailable to the memory subsystem 110 for storing host data.
[0034] After designating the block 210 as isolated, the memory subsystem may perform a stress test on the block 210 to evaluate the health of the block based on the results of the stress test. For example, the memory subsystem 110 may apply a self-heating mechanism to briefly increase the temperature of the memory drive and then observe the behavior of the block 210 caused by this increase in operating temperature. The stress test results of the block may then be evaluated to determine whether the stress test results of the block meet the test criteria.
[0035] In one embodiment, the test criteria may compare the first stress test result of the block 210 with the second stress test result of the healthy block and determine whether the first stress test result is within an expected variance from the second stress test result of the healthy block. In one example, the second stress test result of the healthy block may be retrieved from a predetermined storage location where the test results may be stored for benchmarking purposes. The expected variance may be an expected deviation in percentage points from a benchmark test result of the healthy block.
[0036] At a "failed stress test" operation 248, the memory subsystem 110 may determine that the first stress test result of the block 210 is not within the expected variance of the second stress test result of the healthy block, and therefore may assign the block 210 a retired state 250. In one example, if the threshold voltage of the word line corresponding to the block 210 is found to be greater than 20% higher than the threshold voltage of the second word line corresponding to the healthy block, then the memory subsystem 110 may determine that 20% is not within the expected variance of the healthy test result. The memory subsystem 110 may then determine that the block 210 is unhealthy and may mark the block 210 as retired by assigning it the retired state 250. When the block 210 is in the retired state 250, it may be presumed to be a defective block 260 for the duration of its life, and the memory subsystem 110 may not use the block 210 for storing host data.
[0037] At the "PASSED STRESS TEST" operation 242, if the memory subsystem 110 determines that the first stress test result of the block 210 is within the expected variance of the second stress test result of the healthy block, then the memory subsystem 110 may designate the block 210 as a non-quarantined block by assigning it a healthy state 220, thus allowing the block 210 to exit the quarantine queue. For example, if the threshold voltage of the word line corresponding to the block 210 is found to be slightly higher (5% higher) than the threshold voltage of the second word line corresponding to the healthy block, then the memory subsystem 110 may determine that the block 210 is healthy and may assign the healthy state 220 to the block 210. In another example, if the number of P / E cycles of the block 210 is found to be less than a predetermined threshold of acceptable P / E cycles, then the memory subsystem 110 may similarly determine that the block 210 is healthy and may assign the healthy state 220 to the block 210. The block 210 may then be treated as a normal block by the memory subsystem 110 and may be used to store host data.
[0038] Figure 3 is a flow chart of an example method for managing block retirement for temporary operating conditions according to some embodiments of the present disclosure. Method 300 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, method 300 is performed by Figure 1 The block retirement component 113 is executed. Although shown in a specific order or sequence, the order of the processes can be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. In addition, one or more processes can be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are also possible.
[0039] At operation 310, the processing device detects that a data loss in the block has occurred. Data loss may occur when a read operation of a previously written data bit to the block fails and a subsequent system-level error handling flow fails to recover the data. In one example, the processing device may report data loss when a customer (e.g., when testing the memory subsystem) violates a design specification of the memory subsystem.
[0040] At operation 320, the processing device may identify a behavior criterion that caused data loss in the block to occur. In certain embodiments, the behavior criterion may be whether one of several failure mechanisms satisfies a predetermined threshold criterion (e.g., by having a value exceeding a predetermined threshold). The failure mechanism may include abnormal operating conditions related to the crossover temperature of the block, program / erase cycles (hereinafter P / E cycles), read disturbances, or a combination thereof, as explained in more detail herein above.
[0041] When a behavior criterion has been identified, at operation 330, the processing device may increment a watch list counter associated with the block in response to the occurrence of the behavior criterion. In one embodiment, incrementing the watch list counter may cause the block to be added to a watch list so that the block may be monitored for additional behavior criteria. While on the watch list, the memory block may be monitored for abnormal behavior (e.g., additional data loss events) while the memory block continues to be available to the memory subsystem for storing host data. Whenever a data loss event occurs in the block, the watch list counter for the block may be incremented.
[0042] At operation 340, in response to determining that the value of the watch list counter of the block satisfies the predetermined threshold criteria, the processing device may assign a quarantine status to the block, thereby adding the block to a quarantine queue. For example, a block that appears ten times in the watch list may be added to the quarantine queue. As explained in more detail herein above, the processing device may maintain a quarantine queue to perform certain stress tests on the quarantined block to assess the health of the quarantined block while marking the block as unavailable for storing host data.
[0043] After designating the block as isolated, at operation 350, the processing device may perform a stress test on the block to evaluate the health of the block based on the results of the stress test. The stress test results of the block may then be evaluated to determine whether the stress test results of the block meet the test criteria. In one embodiment, as explained in more detail herein above, the test criteria may compare a first stress test result of the block with a second stress test result of a healthy block and determine whether the first stress test result is within an expected variance from the second stress test result of the healthy block.
[0044] Finally, in response to determining that the stress test results of the block meet the predetermined threshold criteria, the processing device may retire the block by assigning a retirement state to the block at operation 360. For example, if the threshold voltage of the word line corresponding to the block is found to be greater than 20% higher than the threshold voltage of the second word line corresponding to a healthy block, the processing device may determine that the block is unhealthy and may retire the block.
[0045] Figure 4 400 is a flowchart of an example method for adding a block that causes data loss to a watch list before retiring the block according to some embodiments of the present disclosure. The method 400 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 400 is performed by Figure 1 The block retirement component 113 is executed. Although shown in a specific order or sequence, the order of the processes can be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. In addition, one or more processes can be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are also possible.
[0046] At operation 410, the processing device detects the occurrence of data loss in block 210, wherein the data loss is not recoverable by a system-level error handling mechanism, as explained in greater detail herein above. The processing device may determine one or more operating conditions of block 210 at operation 420 to determine whether behavioral criteria associated with the occurrence of data loss in block 210 have existed. In one example, the processing device may determine whether a number of P / E cycles satisfy predetermined criteria, whether a cross-temperature failure has occurred, etc. If the processing device determines at box 425 that at least one operating condition satisfies the predetermined criteria, and therefore determines that at least one failure has occurred, the processing device may mark block 210 for further monitoring (e.g., by adding to a watch list).
[0047] At operation 435, when at least one operating condition is found to satisfy the predetermined criteria, the processing device may increment a watch list counter associated with block 210 in response to the one or more operating conditions satisfying the predetermined criteria. In one embodiment, incrementing the watch list counter may cause the block to be added to a watch list so that the block may be monitored for additional behavior criteria. While on the watch list, the memory block 210 may be monitored for abnormal behavior (e.g., additional data loss events) while continuing to be available to the processing device for use in storing host data. In one example, the watch list counter for block 210 may be incremented each time an operating condition satisfies the predetermined criteria (e.g., when measured at a power-on time of the block).
[0048] If, on the other hand, the operating conditions of block 210 do not satisfy the predetermined criteria, then at operation 450, the processing device may determine that block 210 may operate within normal operating conditions. In certain embodiments, the processing device may reset a watch list counter associated with block 210 to zero, thus ending the monitoring process for block 210 and assigning a healthy state to block 210. In other embodiments, the processing device may determine that a different operating condition of block 210 should be evaluated for monitoring (e.g., crossover temperature, read disturb, P / E cycling, or data retention).
[0049] At operation 445, the processing device may determine whether the value of the watch list counter of block 210 satisfies a configurable threshold criterion. In certain embodiments, if the value of the watch list counter of block 210 exceeds the configurable threshold, then it satisfies the threshold criterion, and vice versa. For example, the processing device may determine that block 210 may be retired after appearing in the watch list significantly less frequently than other blocks within the memory subsystem. In this case, the block may be retired because the memory subsystem may determine that the duplicate data loss event of block 210 is not due to a temporary operating condition applicable to all blocks. If the value of the watch list counter of block 210 satisfies the configurable threshold criterion, then at operation 470, the processing device may retire block 210 by assigning a retirement state to block 210.
[0050] Alternatively, at operation 440, if the processing device determines that the value of the watch list counter of block 210 does not satisfy the configurable threshold criteria, the processing device may continue to monitor the block 210 by keeping the block 210 in the watch list. In one example, the processing device may determine the operating condition of the block 210 at the next reuse of the block 210. The next reuse of the block 210 may be the next power-on event of the memory subsystem 110. The processing device may then proceed to evaluate the operating condition against the predetermined criteria as described above at operation 425.
[0051] Figure 5500 is a flowchart of an example method for marking a block causing data loss as isolated before retiring the block according to some embodiments of the present disclosure. The method 500 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 500 is performed by Figure 1 The block retirement component 113 is executed. Although shown in a specific order or sequence, the order of the processes can be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in a different order, and some processes can be performed in parallel. In addition, one or more processes can be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are also possible.
[0052] At operation 510, the processing device detects a number of data loss occurrences in the block 210, where the data loss occurrences cannot be recovered by the system-level error handling mechanism. In certain embodiments, the data loss occurrences may be caused by one or more failure mechanisms exceeding a predetermined threshold. As explained in more detail herein above, the failure mechanisms may include abnormal operating conditions related to the crossover temperature of the block, P / E cycling, read disturbance, or a combination thereof.
[0053] At operation 520, in response to detecting a data loss event, the block retirement component 113 may mark the block 210 as quarantined by assigning a quarantine status to the block 210. The block retirement component 113 may maintain a quarantine queue to perform certain stress tests on the quarantined blocks to assess the health of the quarantined blocks while marking the blocks as unavailable for storing host data, as explained in more detail herein above.
[0054] After assigning the isolation state to the block 210, at operation 540, the block retirement component 113 may perform a stress test of the block 210 at a high operating temperature. In one example, the block retirement component 113 may apply a self-heating mechanism to briefly increase the temperature of the memory subsystem 110 and then observe the behavior of the block 210 caused by this increase in operating temperature. The stress test results of the block 210 may then be evaluated to determine whether the stress test results of the block meet the test criteria (e.g., by comparing the stress test results to another stress test result of a healthy block).
[0055] At operation 545, the block retirement component 113 may compare the first stress test result of the block 210 with the second stress test result of the healthy block. In one example, the second stress test result of the healthy block may be retrieved from a predetermined storage location where the test results may have been previously stored for benchmarking purposes. The block retirement component 113 then determines, at operation 550, whether the first stress test result is within an acceptable variance from the second stress test result of the healthy block, as explained in more detail herein above. For example, a high operating temperature of the memory subsystem 110 may cause the threshold voltage of the word line corresponding to the block to increase. Therefore, if the threshold voltage of the word line corresponding to the block 210 is found to be greater than 20% higher than the threshold voltage of the second word line corresponding to the healthy block, the block retirement component 113 may determine that 20% is not an acceptable variance and therefore the block 210 may be marked as unusable.
[0056] At operation 555, if the block retirement component 113 determines that the first stress test result of the block 210 is not within an acceptable variance with the second stress test result of the healthy block, the block retirement component 113 may cause the block 210 to be retired by assigning a retirement state to the block 210. As described in more detail herein above, the retired block 210 may not be used to store host data.
[0057] At operation 560, if, on the other hand, the block retirement component 113 determines that the first stress test result of the block 210 is within an acceptable variance with the second stress test result of the healthy block, then the block retirement component 113 may retire the block 210 from the quarantine queue by assigning a healthy state to the block 210. The block 210 may then be used to store host data, as explained in more detail herein above.
[0058] Figure 6 An example machine illustrating a computer system 600 may execute a set of instructions within the computer system 600 for causing the machine to perform any one or more of the methodologies discussed herein. In some embodiments, the computer system 600 may correspond to a computer system that includes, is coupled to, or uses a memory subsystem (e.g., Figure 1 The memory subsystem 110 of the controller may be used to execute the operation of the controller (for example, to execute the operating system to perform the corresponding Figure 1 The operation of the block retirement component 113) of the host system (e.g., Figure 1 In some embodiments, the machine may be connected (e.g., using a network) to other machines in a LAN, an intranet, an extranet, or the Internet. The machine may operate in the capacity of a server or a client machine in a client-server network environment, or in the capacity of a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client machine in a cloud computing infrastructure or environment.
[0059] The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network appliance, a server, a network router, a switch or bridge, an automobile, or any machine capable of executing (sequentially or otherwise) a set of instructions that specify actions to be taken by the machine. In addition, while a single machine is described, the term "machine" should also be construed to include any collection of machines that individually or collectively execute one (or more) sets of instructions to perform any one or more of the methodologies discussed herein.
[0060] The example computer system 600 includes a processing device 602, a main memory 604 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), etc.), a static memory 606 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage system 618, which communicate with each other via a bus 630.
[0061] Processing device 602 represents one or more general processing devices, such as a microprocessor, a central processing unit, or the like. More specifically, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets, or a processor implementing a combination of instruction sets. Processing device 602 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. Processing device 602 is configured to execute instructions 626 for performing the operations and steps discussed herein. Computer system 600 may additionally include a network interface device 608 to communicate on a network 620.
[0062] The data storage system 618 may include a machine-readable storage medium 624 (also referred to as a computer-readable medium) on which is stored one or more sets of instructions 626 or software embodying any one or more of the methods or functions described herein. The instructions 626 may also reside, completely or at least partially, within the main memory 604 and / or within the processing device 602 during execution thereof by the computer system 600, the main memory 604 and the processing device 602 also constituting machine-readable storage media. The machine-readable storage medium 624, the data storage system 618, and / or the main memory 604 may correspond to Figure 1 Memory subsystem 110.
[0063] In one embodiment, instructions 626 include instructions for implementing Figure 1The block retirement component 113 of the embodiment of the present invention is provided with instructions for the functionality of the block retirement component 113. Although the machine-readable storage medium 624 is shown as a single medium in the example embodiment, the term "machine-readable storage medium" should be considered to include a single medium or multiple media storing one or more sets of instructions. The term "machine-readable storage medium" should also be considered to include any medium capable of storing or encoding a set of instructions for execution by a machine and causing the machine to perform any one or more of the methods of the present disclosure. Therefore, the term "machine-readable storage medium" should be considered to include, but not limited to, solid-state memory, optical media, and magnetic media.
[0064] Some portions of the previously detailed description have been presented with respect to algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means by which those skilled in the art of data processing can most effectively communicate the substance of their work to other persons skilled in the art. Algorithms are here and generally considered to be self-consistent sequences of operations that lead to desired results. Operations are operations that require physical control of physical quantities. These quantities are usually, but not necessarily, in the form of electrical or magnetic signals that can be stored, combined, compared, and otherwise manipulated. Sometimes, primarily for general reasons, it has proven convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc.
[0065] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage systems.
[0066] The present disclosure also relates to an apparatus for performing the operations described herein. This apparatus may be specially constructed for the desired purpose, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer readable storage medium, such as, but not limited to, any type of disk (including floppy disks, optical disks, CD-ROMs, and magnetic optical disks), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0067] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general purpose systems may be used with programs according to the teachings herein, or it may prove convenient to construct more specialized devices to perform the methods. The structures of a variety of these systems will be presented as set forth in the description below. In addition, the present disclosure is not described with reference to any particular programming language. It should be appreciated that the teachings of the present disclosure as described herein may be implemented using various programming languages.
[0068] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having stored thereon instructions that can be used to program a computer system (or other electronic device) to perform a process according to the present disclosure. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory ("ROM"), a random access memory ("RAM"), a magnetic disk storage medium, an optical storage medium, a flash memory component, etc.
[0069] In the foregoing description, embodiments of the present disclosure have been described with reference to specific example embodiments thereof. It should be apparent that various modifications may be made to the present disclosure without departing from the broader spirit and scope of embodiments of the present disclosure as set forth in the appended claims. Accordingly, the description and drawings should be viewed in an illustrative rather than a restrictive sense.
Claims
1. A system comprising: Memory components; and A processing device operatively coupled to the memory component to perform the following operations: detecting an occurrence of data loss in a block of the memory component; identifying a behavior criterion associated with an occurrence of said data loss in said block of said memory component; incrementing a counter associated with the block in response to the identified behavior criteria, wherein the value of the counter corresponds to a number of occurrences of a plurality of behavior criteria that caused a data loss in the block to occur, and wherein incrementing the counter associated with the block causes the processing device to monitor the behavior criteria while continuing to designate the block as available to the memory component for storing host data; and In response to determining that the value of the counter satisfies a first threshold criterion: designating the block as an isolated block, performing a first stress test of a plurality of stress tests on the block of the memory component, and In response to the block of the memory component failing the first stress test, retiring the block of the memory component.
2. The system of claim 1, wherein to retire the block, the processing device designates the block as unavailable for use by the memory component to store host data.
3. The system of claim 1, wherein to designate the block as a quarantined block, the processing device designates the block as unavailable for use by the memory component to store host data for a predetermined period of time.
4. The system of claim 1 , wherein in response to the block passing the first stress test, the processing device further performs the following operations: designating the block as a healthy block; and The block is designated as available for use by the memory component to store host data.
5. The system of claim 1, wherein the behavior criteria include a failure mechanism that satisfies a second threshold criterion, the failure mechanism including at least one of a crossover temperature of the block, program / erase cycles, or read disturb. 6 . The system of claim 1 , wherein the plurality of stress tests include at least one of a cross temperature test, a program / erase cycle test, or a read disturb test with respect to the block.
7. The system according to claim 1, wherein the processing device is further configured to determine that the block fails the first stress test. When determining that the block fails the first stress test, the processing device further performs the following operations: retrieving a second stress test result of the healthy block from a predetermined storage location; comparing the first stress test result of the block with the second stress test result of the healthy block; and It is determined that the first stress test result of the block is outside an expected variance from the second stress test result of the healthy block.
8. A method comprising: detecting, by the processing device, that a data loss in a block of the memory component occurs; identifying a behavior criterion associated with an occurrence of said data loss in said block of said memory component; incrementing a counter associated with the block in response to the identified behavior criteria, wherein the value of the counter corresponds to a number of occurrences of a plurality of behavior criteria that caused a data loss in the block to occur, and wherein incrementing the counter associated with the block includes monitoring the behavior criteria while continuing to designate the block as available to the memory component for storing host data; and In response to determining that the value of the counter satisfies a first threshold criterion: designating the block as an isolated block, performing a first stress test of a plurality of stress tests on the block of the memory component, and In response to the block of the memory component failing the first stress test, retiring the block of the memory component.
9. The method of claim 8, wherein retiring the block comprises designating the block as unavailable for use by the memory component to store host data.
10. The method of claim 8, wherein designating the block as a quarantined block comprises designating the block as unavailable to the memory component for storing host data for a predetermined period of time.
11. The method of claim 8, further comprising, in response to the block passing the first stress test: designating the block as a healthy block; and The block is designated as available for use by the memory component to store host data.
12. The method of claim 8, wherein the behavior criteria include a failure mechanism that satisfies a second threshold criterion, the failure mechanism including at least one of a crossover temperature of the block, program / erase cycles, or read disturb. 13 . The method of claim 8 , wherein the plurality of stress tests include at least one of a cross temperature test, a program / erase cycle test, or a read disturb test with respect to the block.
14. The method of claim 8, further comprising determining that the block has failed the first stress test, wherein determining that the block has failed the first stress test comprises: retrieving a second stress test result of the healthy block from a predetermined storage location; comparing the first stress test result of the block with the second stress test result of the healthy block; and It is determined that the first stress test result of the block is outside an expected variance from the second stress test result of the healthy block.
15. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to: identifying a behavior criterion associated with an occurrence of data loss in a block of a memory component; incrementing a counter associated with the block in response to the identified behavior criteria, wherein the value of the counter corresponds to a number of occurrences of a plurality of behavior criteria that caused a data loss in the block to occur, and wherein incrementing the counter associated with the block includes monitoring the behavior criteria while continuing to designate the block as available to the memory component for storing host data; and In response to determining that the value of the counter satisfies a first threshold criterion: designating the block as an isolated block, performing a first stress test of a plurality of stress tests on the block of the memory component, and In response to the block of the memory component failing the first stress test, retiring the block of the memory component.
16. The non-transitory computer-readable storage medium of claim 15, wherein to determine that the block fails the first stress test, the processing device further performs the following operations: retrieving a second stress test result of the healthy block from a predetermined storage location; comparing the first stress test result of the block with the second stress test result of the healthy block; and It is determined that the first stress test result of the block is outside an expected variance from the second stress test result of the healthy block.
17. The non-transitory computer-readable storage medium of claim 15, wherein the behavioral criteria include a failure mechanism that satisfies a second threshold criterion, the failure mechanism including at least one of a cross temperature, a program / erase cycle, or a read disturb of the block, and wherein the plurality of stress tests of the block include at least one of a cross temperature test, a program / erase cycle test, or a read disturb test on the block.
18. The non-transitory computer-readable storage medium of claim 15, wherein in response to determining that the block fails the first stress test, the processing device further performs the following operations: designating the block as a healthy block; and The block is designated as available for use by the memory component to store host data.
Citation Information
Patent Citations
Devices, systems, and methods for predicting faults in solid-state storage devices.
CN102272731A
Test method for nonvolatile memory
US20140289559A1