DATA STORAGE DEVICE AND METHOD FOR ADVANCED RECOVERY BY HARDWARE RESET OF ONE OF ITS DISCRETE COMPONENTS
A hardware reset mechanism for non-responsive memory dies in data storage devices addresses the inefficiencies of conventional recovery methods, enhancing recovery efficiency and optimizing performance and capacity by minimizing the number of non-functional dies.
Patent Information
- Application Number
- DE112023003545
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-07-18
- Filing Date
- 2023-12-19
- Publication Date
- 2025-08-07
AI Technical Summary
Existing data storage devices face challenges in efficiently recovering from non-responsive memory dies without causing performance degradation and over-provisioning, particularly in high-capacity systems where conventional methods like XOR recovery can lead to data loss and reduced efficiency.
Implementing a hardware reset mechanism for non-responsive memory dies, coordinated with the host, to perform a controlled power-off and power-on of the affected components, allowing for extended recovery without requiring host intervention or queue clearing, thus reducing the number of dies taken out of service and improving performance and capacity.
The hardware reset approach enhances recovery efficiency, minimizes performance degradation, and optimizes resource utilization by reducing the number of non-functional dies, thereby improving the overall performance and capacity of data storage devices.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of the entire contents of U.S. Non-Provisional Application No. 18 / 223,122, entitled "Data Storage Device and Method for Enhanced Recovery Through a Hardware Reset of One of Its Discrete Components," filed on July 18, 2023, in the U.S. Patent and Trademark Office, which claims priority to U.S. Provisional Application No. 63 / 449,770, filed on March 3, 2023, and is hereby incorporated by reference for all purposes. BACKGROUND
[0002] With the growth of storage usage in enterprises and the cloud, the need for data correction techniques and highly reliable data storage devices is also increasing. Some data storage device controllers use various data protection mechanisms to help ensure a low read error rate and ensure that the data returned to a host is free of integrity errors, such as the minimum bit error rate (UBER) or mean time between failures (MTBF). BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1A is a block diagram of a data storage device of one embodiment. Fig. 1B is a block diagram illustrating a storage module of one embodiment. Fig. Figure 1C is a block diagram illustrating a hierarchical storage system of one embodiment. Fig. Figure 2A is a block diagram showing components of the controller of the Fig. 1A according to one embodiment. Fig. Figure 2B is a block diagram showing components of the Fig. 1A illustrates a memory data storage device according to one embodiment. Fig. 3 is a block diagram of a host and a data storage device of one embodiment. Fig. 4 is a flowchart of a method of one embodiment for die recovery. Fig. 5 is a flowchart of a method of one embodiment for an enhanced recovery operation. Fig. 6 is a flowchart of a method of one embodiment for an enhanced recovery operation. Fig. 7 is a flowchart of an embodiment method for an enhanced recovery operation based on various components. DETAILED DESCRIPTIONOverview
[0003] By way of introduction, the following embodiments relate to a data storage device and a method for enhanced recovery through a hardware reset of one of its discrete components. In one embodiment, a data storage device is provided, comprising: a plurality of memory dies and a controller. The controller is configured to: in response to determining that a subset of the plurality of memory dies is unresponsive, inform a host that the data storage device is performing a hardware reset on the subset of the plurality of memory dies; in response to receiving an acknowledgment from the host, perform the hardware reset on the subset of the plurality of memory dies; and inform the host after the hardware reset has been performed on the subset of the plurality of memory dies.
[0004] In some embodiments, the hardware reset is initiated by the data storage device and not by the host.
[0005] In some embodiments, a hardware reset is performed on the subset of the plurality of memory dies without clearing the host input / output queue(s) for the subset of the plurality of memory dies.
[0006] In some embodiments, the controller is further configured to receive, from the host, a memory access command redirected from the subset of the plurality of memory dies to another memory die of the plurality of memory dies.
[0007] In some embodiments, the controller is further configured to perform the hardware reset on the subset of the plurality of memory dies by: sending a command to all memory dies of the plurality of memory dies to ignore a hardware reset command, wherein, since the subset of the plurality of memory dies is unresponsive, the subset of the plurality of memory dies does not receive the command to ignore the hardware reset command; and sending the hardware reset command on a communication channel shared by the plurality of memory dies.
[0008] In some embodiments, the controller is further configured to: determine whether the subset of the plurality of memory dies has been determined to be unresponsive more than a threshold number of times over a period of time; in response to determining that the subset of the plurality of memory dies has not been determined to be unresponsive more than the threshold number of times over the period of time, performing the hardware reset on the subset of the plurality of memory dies; and, in response to determining that the subset of the plurality of memory dies has been determined to be unresponsive more than the threshold number of times over the period of time, decommissioning the subset of the plurality of memory dies.
[0009] In some embodiments, the controller is further configured to inform the host that the data storage device performs the hardware reset in response to determining that the host has enabled an extended recovery operation.
[0010] In some embodiments, the subset of the plurality of memory dies comprises a single memory die.
[0011] In some embodiments, the subset of the plurality of memory dies comprises more than one memory die.
[0012] In some embodiments, the plurality of memory dies comprise three-dimensional memory dies.
[0013] In another embodiment, a method is provided that is performed in a data storage device that communicates with a host. The method includes: determining that a component in the data storage device requires a hardware reset; informing the host that the data storage device is performing a hardware reset of the component; performing the hardware reset of the component; and informing the host after the hardware reset of the component has been performed.
[0014] In some embodiments, the component comprises a memory die.
[0015] In some embodiments, the component comprises a portion of a controller of the data storage device.
[0016] In some embodiments, the component comprises one or more of: a dynamic random access memory (DRAM), an electrically erasable programmable read-only memory (EEPROM), a component in an application-specific integrated circuit (ASIC), a capacitor, and a bus controller.
[0017] In some embodiments, in response to determining that the component is unresponsive, the data storage device determines that the component requires a hardware reset.
[0018] In some embodiments, in response to determining that the component is experiencing irregular power consumption, the data storage device determines that the component requires a hardware reset.
[0019] In some embodiments, in response to determining that the component has a certain read and / or write latency profile, the data storage device determines that the component requires a hardware reset.
[0020] In some embodiments, the method further comprises receiving an instruction from the host to enable functionality in the data storage device to perform a hardware reset of the component.
[0021] In some embodiments, a hardware reset of the component is performed by the data storage device and not by the host.
[0022] In another embodiment, a data storage device is provided, comprising: a plurality of memory dies; means for determining that one or more memory dies of the plurality of memory dies are inoperative; means for informing a host communicating with the data storage device that an extended rebuild operation is being performed on the one or more memory dies such that a response to a host command may be delayed; and means for performing the extended rebuild operation on the one or more memory dies.
[0023] Other embodiments are possible, and each of the embodiments may be used alone or in combination with each other. Accordingly, various embodiments will now be described with reference to the accompanying drawings. Embodiments
[0024] The following embodiments relate to a data storage device (DSD). As used herein, "data storage device" refers to a device that stores data. Examples of DSDs include, but are not limited to, hard disk drives (HDDs), solid-state drives (SSDs), tape drives, hybrid drives, etc. Details of example DSDs are provided below.
[0025] Data storage devices suitable for use in implementing aspects of these embodiments are described in the Fig. 1A to 1C. Fig. 1A is a block diagram illustrating a data storage device 100 according to an embodiment of the subject matter described herein. Referring to Fig. 1A, the data storage device 100 includes a controller 102 and a non-volatile memory, which may consist of one or more non-volatile memory dies 104. As used herein, the term "die" refers to the collection of non-volatile memory cells and associated circuitry for managing the physical operation of these non-volatile memory cells formed on a single semiconductor substrate. The controller 102 is connected to a host system and transmits command sequences for read, program, and erase operations to the non-volatile memory die 104.
[0026] The controller 102 (which may be a non-volatile memory controller (e.g., a flash memory, resistive random access memory (ReRAM), phase change memory (PCM), or magnetoresistive random access memory (MRAM) controller)) may, for example, take the form of processing circuitry, a microprocessor or processor, and a computer-readable medium storing computer-readable program code (e.g., firmware) executable by the (micro)processor, logic gates, switches, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. The controller 102 may be configured with hardware and / or firmware to perform the various functions described below and illustrated in the flowcharts.In addition, some of the components shown as internal to the controller may also be stored external to the controller, and other components may be used. Furthermore, the term "operatively associated with" may mean a direct connection to, or an indirect (wired or wireless) connection to, one or more components that may or may not be shown or described herein.
[0027] As used herein, a non-volatile memory controller is a device that manages data stored in non-volatile memory and communicates with a host, such as a computer or electronic device. A non-volatile memory controller may have various functions in addition to the specific functions described here. For example, the non-volatile memory controller may format the non-volatile memory to ensure that the memory functions properly, locate faulty non-volatile memory cells, and allocate spare cells to replace future failed cells. A portion of the spare cells may be used to store firmware to operate the non-volatile memory controller and implement other functions.During operation, when a host needs to read or write data to non-volatile memory, it can communicate with the non-volatile memory controller. If the host provides a logical address to which data should be read / written, the non-volatile memory controller can convert the logical address received from the host into a physical address in the non-volatile memory. (Alternatively, the host can provide the physical address.) The non-volatile memory controller can also perform various memory management functions, such as wear leveling (spreading out write operations to avoid wearing out certain memory blocks that would otherwise be repeatedly written to) and automatic garbage collection (when a block is full, only the valid data pages are moved to a new block, allowing the full block to be erased and reused).
[0028] The non-volatile memory die 104 may include any suitable non-volatile storage medium, including resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), phase-change memory (PCM), NAND flash memory cells, and / or NOR flash memory cells. The memory cells may be in the form of solid-state memory cells (e.g., flash memory cells) and may be programmable once, multiple times, or many times. The memory cells may also be single-level cells (SLC), multiple-level cells (MLC) (e.g., dual-level cells, triple-level cells (TLC), quad-level cells (QLC), etc.), or utilize technologies with other memory cell levels now known or later developed. Furthermore, the memory cells may be fabricated two-dimensionally or three-dimensionally.
[0029] The interface between controller 102 and non-volatile memory die 104 can be any suitable flash interface, such as Toggle Mode 200, 400, or 800. In one embodiment, data storage device 100 can be a card-based system, such as a Secure Digital (SD) card or a Micro Secure Digital (Micro SD) card. In an alternative embodiment, data storage device 100 can be part of an embedded data storage device.
[0030] Although in the Fig. 1A, the data storage device 100 (sometimes referred to herein as a storage module) includes a single channel between the controller 102 and the non-volatile memory die 104, the subject matter described herein is not limited to having a single memory channel. For example, in some architectures (such as those shown in Fig. 1B and Fig. 1C), depending on the controller functions, two, four, eight, or more memory channels may be present between the controller and the memory device. In all embodiments described herein, more than a single channel may be present between the controller and the memory die, even though only a single channel is shown in the drawings.
[0031] Fig. 1B illustrates a storage module 200 including a plurality of non-volatile data storage devices 100. As such, the storage module 200 may include a storage controller 202 connected to a host and to the data storage device 204, which includes a plurality of data storage devices 100. The interface between the storage controller 202 and the data storage devices 100 may be a bus interface, such as a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, or a Double Data Rate (DDR) interface. The storage module 200, in one embodiment, may be a solid-state drive (SSD) or a non-volatile dual in-line storage module (NVDIMM) such as those found in server PCs or portable computing devices such as laptops and tablet computers.
[0032] Fig. Figure 1C is a block diagram illustrating a hierarchical storage system. A hierarchical storage system 250 includes a plurality of storage controllers 202, each of which controls a corresponding data storage device 204. Host systems 252 can access storage within the storage system 250 via a bus interface. In one embodiment, the bus interface can be a Non-Volatile Memory Express (NVMe) interface or a Fibre Channel over Ethernet (FCoE) interface. In one embodiment, the Fig. The system illustrated in Figure 1C may be a rack-mountable mass storage system accessible by multiple host computers, such as might be found in a data center or other location where mass storage is needed.
[0033] Fig. Figure 2A is a block diagram illustrating the components of controller 102 in more detail. Controller 102 includes a front-end module 108 connected to a host, a back-end module 110 connected to one or more non-volatile memory dies 104, and various other modules that perform functions that will now be described in detail. A module may take the form of a bundled functional hardware unit designed for use with other components, a piece of program code (e.g., software or firmware) that can be executed by a (micro)processor or processing logic that typically performs a specific function of related functions, or a self-contained hardware or software component that is connected, for example, to a larger system.In addition, "means" for performing a function may be implemented with at least one of the structures specified herein for the controller and may be pure hardware or a combination of hardware and computer-readable program code.
[0034] Referring again to the modules of controller 102, a buffer manager / bus controller 114 manages buffers in random access memory (RAM) 116 and controls the internal bus arbitration of controller 102. A read-only memory (ROM) 118 stores the system startup code. Although in Fig. 2A as being located separately from the controller 102, in other embodiments, one or both of the RAM 116 and the ROM 118 may be located within the controller. In still other embodiments, portions of the RAM and ROM may be located both within the controller 102 and external to the controller.
[0035] The front-end module 108 includes a host interface 120 and a physical layer interface (PHY) 122, which provide the electrical interface with the host or the next-level storage controller. The choice of host interface 120 type may depend on the storage type used. Examples of host interfaces 120 include, but are not limited to, SATA, SATA Express, Serially Attached Small Computer System Interface (SAS), Fibre Channel, Universal Serial Bus (USB), PCIe, and NVMe. The host interface 120 typically facilitates the transfer of data, control signals, and timing signals.
[0036] The back-end module 110 includes an error-correcting code (ECC) engine 124, which encodes the data bytes received from the host and decodes and corrects errors from the data bytes read from the non-volatile memory. A command sequencer 126 generates command sequences, such as program and erase command sequences, to be transmitted to the non-volatile memory die 104. A RAID (Redundant Array of Independent Drives) module 128 manages the generation of RAID parity and the recovery of corrupted data. RAID parity can be used as an additional level of integrity protection for the data being written to the storage device 104. In some cases, the RAID module 128 can be part of the ECC engine 124. A memory interface 130 provides the command sequences to the non-volatile memory die 104 and receives status information from the non-volatile memory die 104.In one embodiment, memory interface 130 may be a double data rate (DDR) interface, such as a Toggle Mode 200, 400, or 800 interface. A flash control layer 132 controls the overall operation of back-end module 110.
[0037] The data storage device 100 also includes other separate components 140, such as external electrical interfaces, external RAM, resistors, capacitors, or other components that may be connected to the controller 102. In alternative embodiments, one or more of the physical layer interface 122, the RAID module 128, the media management layer 138, and the buffer management / bus controller 114 are optional components that are not required in the controller 102.
[0038] Fig. Figure 2B is a block diagram illustrating the components of the non-volatile memory die 104 in more detail. The non-volatile memory die 104 includes peripheral circuitry 141 and a non-volatile memory array 142. The non-volatile memory array 142 includes the non-volatile memory cells used to store data. The non-volatile memory cells may be any suitable non-volatile memory cells, including ReRAM, MRAM, PCM, NAND flash memory cells, and / or NOR flash memory cells in a two-dimensional and / or three-dimensional configuration.
[0039] The non-volatile memory die 104 further includes a data cache 156 that temporarily stores data. The peripheral circuitry 141 includes a state machine 152 that provides status information to the controller 102.
[0040] With reference again to Fig. 2A, the flash control layer 132 (referred to herein as the flash translation layer (FTL) or more generally as the "media management layer" because the memory may not be flash) handles flash errors and communicates with the host. In particular, the FTL, which may be an algorithm in the firmware, is responsible for the internal operations of memory management, translating writes from the host into writes to memory 104. The FTL may be necessary because memory 104 may have a limited lifetime, can only be written to in multiples of pages, and / or cannot be written to unless erased as a block. The FTL understands these potential limitations of memory 104, which may not be visible to the host. Accordingly, the FTL attempts to translate the writes from the host into writes to memory 104.
[0041] The FTL may include a logical-to-physical address (L2P) mapping (sometimes referred to herein as a table or data structure) and allocated cache memory. In this way, the FTL translates logical block addresses ("LBAs") from the host into physical addresses in memory 104. The FTL may include other features, such as, but not limited to, power-off recovery (so that the FTL data structures can be recovered in the event of a sudden loss of power) and wear-leveling (so that wear is even across memory blocks to prevent certain blocks from becoming excessively worn, which would lead to a greater probability of failure).
[0042] Referring again to the drawings, Fig. 3 is a block diagram of a host 300 and a data storage device 100 of one embodiment. The host 300 may take any suitable form, including, but not limited to, a computer, a mobile phone, a tablet, a wearable device, a digital video recorder, a surveillance system, etc. The host 300 in this embodiment (herein, a computing device) includes a processor 330 and a memory 340. In one embodiment, computer-readable program code stored in the host memory 340 configures the host processor 330 to perform the actions described herein. Therefore, actions performed by the host 300 are sometimes considered herein to be performed by an application (computer-readable program code) executing on the host 300.For example, host 300 may be configured to send data (e.g., initially stored in host memory 340) to data storage device 100 for storage in memory 104 of the data storage device.
[0043] As mentioned above, a data storage device controller can employ various data protection mechanisms to ensure a low read error rate and ensure that data returned to the host is free of integrity errors. One such common mechanism in a flash data storage device is exclusive-or (XOR) protection for block- or die-level recovery, which is used in high-capacity products that require a low unrecoverable bit error rate (UBER) or mean time between failures (MTBF). These elements can include error detection and correction on storage / memory elements located both in the data storage device controller and on the non-volatile memory dies.Connection problems between the controller and the storage dies can cause integrity errors that are monitored and detected in the controller.
[0044] When errors are detected in memory, a read error often occurs when the data is considered unrecoverable. If a die in server-based flash storage devices becomes unresponsive, it can be removed from the system, and an attempt can be made to recover the data using the XOR mechanism, but this may result in a loss of performance during the recovery process. This process is described in flowchart 400 in Fig. 4. As shown in Fig. As shown in Figure 4, the non-response of a die is detected (action 410). A die removal process is then initiated, in which data is recovered using an exclusive-or (XOR) operation and written to another die (action 420). The end result is that the die is removed without any data being written to or read from it (action 430). As mentioned above, the host may experience performance degradation during the die removal process due to the significant number of read recovery operations. Additionally, depending on the number of removed dies, reduced overprovisioning (OP) and, consequently, performance degradation may also occur.
[0045] This recovery scheme can only be applied to very high-capacity points that can afford the loss of a die. The reason for taking the die out of service when it becomes unresponsive is to enable hot recovery with minimal impact on the quality of service (QoS) for ongoing input / output (IO) operations. Furthermore, in some cases, there is no way to reestablish communication with the die. In many cases, however, the reason the die becomes unresponsive is due to temporary corruption of its internal state machine, which can be resolved by a power cycle (i.e., cycling the memory power), and then the component can function correctly without any additional risk of data loss.
[0046] The following embodiments can be used to enhance the recovery of a memory element with an internal reset / power cycle of a non-functional memory die through a protocol agreed upon with the host. This can significantly reduce the likelihood of negative effects of die removal. For example, performing a hardware reset (e.g., power cycling or performing a hard reset) of a memory die can result in the erasure of information stored in the memory die's latches / cache, state machine information, interface settings / information, etc., which may be corrupted and lead to memory die failure. This allows data storage device components (e.g., NAND dies) to be recovered, which is not possible with techniques such as exclusive-or (XOR) or a software reset.This can also help reduce the number of decommissioned dies and improve performance, overprovisioning, and capacity.
[0047] In one embodiment, if a system element (e.g., a storage die) becomes unresponsive, the controller 102 in the data storage device 100 may attempt to reset it through a series of steps after reaching an agreement with the host 300. If a storage die 104 becomes unresponsive, the controller 102 of the data storage device 100 may indicate its recovery attempt (e.g., by warning of a brief IO stall) through a handshake with the host 300, but avoid decommissioning the affected die 104. This does not require the host 300 to perform a power cycle or reinitialize the NVMe configuration. It also does not require the host 300 to clean up the drive's queued IOs.The host 300 can authorize the use of advanced recovery for each of the supported components (it can choose which components it would support advanced recovery for).
[0048] This embodiment is shown in flowchart 500 of Fig. 5. As in Fig. 5, after detecting a non-response from the die (action 510), the controller 102 of the data storage device 100 indicates to the host 300 that an extended restore is required (action 520). In response, the host 300 stops transmitting operations to the storage die or accepts a greater latency and notifies the data storage device 100 that it can begin the extended restore process (action 530). The controller 102 of the data storage device 100 then executes an extended restore by performing a hardware reset of the storage die 104 and notifying the host 300 of the completion of the execution (action 540). The host 300 then resumes operation of the respective data storage device 100 (action 550).
[0049] As used herein, "hardware reset" can refer to a power cycle or hard reset that returns a component to the state it was in when it left the factory. A hardware reset can remove settings, applications, and user data. In contrast, a "firmware (or soft) reset" can refer to a reboot of a component to erase data from volatile memory and restart an application without completely shutting down the component.
[0050] Thus, in this embodiment, a handshake between host 300 and controller 102 occurs when controller 102 first notifies host 300 of the need to perform an extended restore and that a particular memory die (or dies) is unavailable (i.e., that host 300 should avoid sending memory access commands to the unresponsive memory die(s) or accept greater latency). In response, host 300 may stop all memory access (IO) operations (e.g., read and / or write commands) to the memory die(s) that are undergoing a data storage device-controlled reset. For example, host 300 may prepare for an IO stall (e.g., byThe host 300 will wait for pending IOs to complete and redirect the next IOs to other storage dies) and respond when it is ready for the controller 102 to perform the extended rebuild operation to reset the storage die. The host 300 does not need to perform a power cycle or reinitialize the NVMe configuration because the controller 102 can perform the power cycle of the storage die itself. The host 300 also does not need to clear the queues. Once the extended rebuild is complete, the host 300 can proceed as usual.
[0051] Additionally, in some memory architectures, such as NAND, a reset pin (or other type of communication channel) is shared by all memory dies (e.g., where a pin-per-die or address-based system may not be feasible given the large number of dies in the device). Because only a subset of memory dies (e.g., one) may need to be reset, a command can be sent to the other memory dies to ignore the reset pin. This allows controller 102 to direct a reset only to a specific memory die(s). The nonfunctional die(s) cannot receive commands, so the "ignore" command is received only by the functional memory dies.
[0052] In another embodiment, there may be a heuristic that the previous decommissioning protocol is triggered if a die shows "no response" multiple times within a certain period of time. This embodiment is illustrated in flowchart 600 in Fig. 6. As shown in Fig. 6, after detecting the non-responsiveness of a die (action 610), the controller 102 determines whether the die has shown no response for a period of time (e.g., in the last X minutes) (action 620). (Action 620 may be modified with a different heuristic that also considers other irregular behavior of the die in question besides "no response," such as irregular power supply, read / write latency profile, or any other characteristic indicating that a component in the data storage device 100 is malfunctioning or experiencing a fault.) If the die has shown no non-responsiveness for the period of time, the controller 102 performs the extended recovery operation with a host handshake (e.g., as in Fig. 5) (action 630). However, if the die has shown a non-response over the period, the controller 102 takes the die out of service (action 640).
[0053] In another embodiment, other elements within storage controller 102 may trigger an extended recovery operation. For example, dynamic random access memory ("DRAM"), electrically erasable programmable read-only memory ("EEPROM"), or portions of the ASIC itself may be cycled (or reset) to recover from a memory- or interface-related transient error. Extended recovery may also be applied to non-memory die / storage elements, such as, but not limited to, capacitors or System Management Bus (SMBus) controllers. Further, extended recovery may be enabled or disabled separately by host 300 for each of these components, possibly based on the expected IO quiescent duration reported by data storage device 100.The host 300 can also instruct the drive to perform the rebuild internally without additional notification to the host 300.
[0054] If storage controller 102 identifies a problem with a recoverable component, it can check whether host 300 has enabled extended recovery for the corresponding component. Depending on the component, it can initiate the recovery operation. For example, if DRAM is unavailable during the specified time period, controller 102 can still access local storage if some urgent host operations need to be completed.
[0055] Fig. Figure 7 shows a flowchart 700 illustrating this embodiment. As in Fig. 7, the host 300 configures (enables / disables) enhanced recovery for different data storage device components (action 710). A component failure is detected for a component enabled for enhanced recovery (action 720). The controller 102 of the data storage device 100 then indicates to the host 300 that enhanced recovery is required for the corresponding component (action 730). The host 300 operates according to the failed component and notifies the data storage device 100 that it can begin the enhanced recovery process (action 740). The controller 102 of the data storage device 100 then executes enhanced recovery and notifies the host 300 upon completion of execution (action 750). The host 300 then resumes normal operation of the data storage device 100 (action 760).
[0056] In this embodiment, host 300 configures the components and possibly the corresponding allowable recovery time, which also allows for determining whether controller 102 initiates the extended recovery handshake depending on the type of error. If a fault is then detected in a component that host 300 has approved for extended recovery, controller 102 may initiate the handshake process with host 300. The response actions from host 300 may depend on the type of faulty component in action 740 in Fig. 7 be dependent.
[0057] Several advantages are associated with these embodiments. For example, the embodiments can reduce the number of decommissioned dies, thereby improving performance, overprovisioning, and capacity, thereby enhancing our storage devices in high-end server products.
[0058] Finally, as mentioned above, any suitable memory type may be used. Semiconductor memory devices include volatile memory devices such as dynamic random access memory ("DRAM") or static random access memory ("SRAM"), non-volatile memory devices such as resistive random access memory ("ReRAM"), electrically erasable programmable read-only memory ("EEPROM"), flash memory (which can also be considered a subset of EEPROM), ferroelectric random access memory ("FRAM"), and magnetoresistive random access memory ("MRAM"), as well as other semiconductor elements capable of storing information. Each memory device type can have different configurations. For example, flash memory devices can be configured in a NAND or NOR configuration.
[0059] The memory devices can be formed from passive and / or active elements in any combination. As a non-limiting example, passive semiconductor memory elements include ReRAM device elements, which in some embodiments include a resistive switching storage element, such as an antifuse, a phase-change material, etc., and optionally a steering element, such as a diode, etc. As another non-limiting example, active semiconductor memory elements include EEPROM and Flash memory device elements, which in some embodiments include elements containing a charge storage region, such as a floating gate, conductive nanoparticles, or a dielectric charge storage material.
[0060] Multiple memory elements can be configured to be connected in series or so that each element is individually accessible. As a non-limiting example, flash memory devices in a NAND (NAND memory) configuration typically include memory elements connected in series. A NAND memory array can be configured so that the array consists of multiple memory strings, where a string consists of multiple memory elements that share a single bitline and are accessed as a group. Alternatively, memory elements can be configured so that each element can be accessed individually, such as a NOR memory array. NAND and NOR memory configurations are examples, and memory elements can be configured differently.
[0061] The semiconductor memory elements located within and / or above a substrate can be arranged in two or three dimensions, for example as a two-dimensional memory structure or as a three-dimensional memory structure.
[0062] In a two-dimensional memory structure, the semiconductor memory elements are arranged in a single plane or a single memory device plane. Typically, in a two-dimensional memory structure, memory elements are arranged in a plane (e.g., a plane in the xz direction) that is substantially parallel to a major surface of a substrate supporting the memory elements. The substrate may be a wafer over or within which the layer of memory elements is formed, or it may be a carrier substrate that is attached to the memory elements after they have been formed. As a non-limiting example, the substrate may include a semiconductor such as silicon.
[0063] The memory elements may be arranged in an ordered array, such as in a plurality of rows and / or columns, within the single memory device level. However, the memory elements may be arranged in non-regular or non-orthogonal configurations. The memory elements may each have two or more electrodes or contact lines, such as bit lines and word lines.
[0064] A three-dimensional memory array is arranged so that memory elements occupy multiple levels or multiple memory device levels, thereby forming a structure in three dimensions (i.e., in the x, y, and z directions, with the y direction being substantially perpendicular and the x and z directions being substantially parallel to the main surface of the substrate).
[0065] As a non-limiting example, a three-dimensional memory structure may be arranged vertically as a stack of multiple two-dimensional memory device levels. As a further non-limiting example, a three-dimensional memory array may be arranged as multiple vertical columns (e.g., columns extending substantially perpendicular to the main surface of the substrate, i.e., in the y-direction), with each column having multiple memory elements in each column. The columns may be arranged in a two-dimensional configuration, e.g., in an xz-plane, resulting in a three-dimensional array of memory elements with elements on multiple vertically stacked memory levels. Other configurations of memory elements in three dimensions may also form a three-dimensional memory array.
[0066] As a non-limiting example, in a three-dimensional NAND memory array, the storage elements may be coupled together to form a NAND string within a single horizontal (e.g., xz) memory device plane. Alternatively, the storage elements may be coupled together to form a vertical NAND string spanning multiple horizontal memory device planes. Other three-dimensional configurations are conceivable, where some NAND strings contain storage elements in a single memory plane, while other strings contain storage elements spanning multiple memory planes. Three-dimensional memory arrays may also be designed in a NOR configuration and in a ReRAM configuration.
[0067] Typically, in a monolithic three-dimensional memory array, one or more memory device levels are formed over a single substrate. Optionally, the monolithic three-dimensional memory array may also include one or more memory layers located at least partially within the single substrate. As a non-limiting example, the substrate may include a semiconductor such as silicon. In a monolithic three-dimensional array, the layers forming each memory device level of the array are typically formed on top of the layers of the array's underlying memory device levels. However, layers of adjacent memory device levels of a monolithic three-dimensional memory array may be shared or may have intermediate layers between the memory device levels.
[0068] On the other hand, two-dimensional arrays can be formed separately and then stacked together to form a non-monolithic memory device with multiple memory layers. For example, non-monolithic stacked memories can be constructed by forming memory planes on separate substrates and then stacking the memory planes on top of each other. The substrates can be thinned or removed from the memory planes before stacking, but because the memory planes are first formed over separate substrates, the resulting memory arrays are not monolithic three-dimensional memory arrays. Furthermore, multiple two-dimensional memory arrays or three-dimensional memory arrays (monolithic or non-monolithic) can be formed on separate chips and then bundled to form a stacked-chip memory device.
[0069] Operating and communicating with the memory elements typically requires associated circuitry. As non-limiting examples, memory devices may include circuitry used to control and drive memory elements to perform functions such as programming and reading. This associated circuitry may be located on the same substrate as the memory elements and / or on a separate substrate. For example, a controller for memory read / write operations may be located on a separate controller chip and / or on the same substrate as the memory elements.
[0070] Those skilled in the art will recognize that this invention is not limited to the described two-dimensional and three-dimensional structures, but covers all relevant memory structures within the spirit and scope of the invention as described herein and understood by those skilled in the art.
[0071] The foregoing detailed description is intended to illustrate selected forms the invention may take, and not as a definition of the invention. Only the following claims, including all equivalents, are intended to define the scope of the claimed invention. Finally, it is to be understood that any aspect of any embodiment described herein may be used alone or in combination with one another. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] US 18 / 223,122
[0001] US 63 / 449,770
[0001]
Claims
[1] Data storage device comprising: a variety of memory dies and a controller configured to communicate with the plurality of memory dies and further configured to: in response to determining that a subset of the plurality of memory dies is unresponsive, informing a host that the data storage device is performing a hardware reset on the subset of the plurality of memory dies; in response to receiving an acknowledgment from the host, performing the hardware reset on the subset of the plurality of memory dies and Informing the host after the hardware reset has been performed on the subset of the plurality of memory dies. [2] The data storage device of claim 1, wherein the hardware reset is initiated by the data storage device and not by the host. [3] The data storage device of claim 1, wherein the subset of the plurality of memory dies undergoes a hardware reset without clearing the host input / output queue(s) for the subset of the plurality of memory dies. [4] The data storage device of claim 1, wherein the controller is further configured to receive, from the host, a memory access command redirected from the subset of the plurality of memory dies to another memory die of the plurality of memory dies. [5] The data storage device of claim 1, wherein the controller is further configured to perform the hardware reset on the subset of the plurality of memory dies by: Sending a command to all memory dies of the plurality of memory dies to ignore a hardware reset command, wherein, since the subset of the plurality of memory dies does not respond, the subset of the plurality of memory dies does not receive the command to ignore the hardware reset command; and Sending the hardware reset command on a communication channel shared by the plurality of memory dies. [6] The data storage device of claim 1, wherein the controller is further configured to: Determining whether the subset of the plurality of memory dies was found to be unresponsive more than a threshold number of times over a period of time; in response to determining that the subset of the plurality of memory dies was not determined to be unresponsive more than the threshold number of times over the period, performing the hardware reset on the subset of the plurality of memory dies; and in response to determining that the subset of the plurality of memory dies was found to be unresponsive more than the threshold number of times over the period, decommissioning the subset of the plurality of memory dies. [7] The data storage device of claim 1, wherein the controller is further configured to inform the host that the data storage device performs the hardware reset in response to determining that the host has enabled an extended recovery operation. [8] The data storage device of claim 1, wherein the subset of the plurality of memory dies comprises a single memory die. [9] The data storage device of claim 1, wherein the subset of the plurality of memory dies comprises more than one memory die. [10] The data storage device of claim 1, wherein the plurality of memory dies comprise three-dimensional memory dies. [11] Method comprising: Performing the following in a data storage device in communication with a host: Determining that a component in the data storage device requires a hardware reset; Informing the host that the data storage device is performing a hardware reset of the component; Performing a hardware reset of the component and Inform the host after the component hardware reset has been performed. [12] The method of claim 11, wherein the component comprises a memory die. [13] The method of claim 11, wherein the component comprises part of a controller of the data storage device. [14] The method of claim 11, wherein the component comprises one or more of: a dynamic random access memory (DRAM), an electrically erasable programmable read-only memory (EEPROM), a component in an application-specific integrated circuit (ASIC), a capacitor, and a bus controller. [15] The method of claim 11, wherein the data storage device, in response to determining that the component is unresponsive, determines that the component requires a hardware reset. [16] The method of claim 11, wherein the data storage device, in response to determining that the component has irregular power consumption, determines that the component requires a hardware reset. [17] The method of claim 11, wherein the data storage device, in response to determining that the component has a certain read and / or write latency profile, determines that the component requires a hardware reset. [18] The method of claim 11, further comprising receiving an instruction from the host to enable functionality in the data storage device to perform a hardware reset of the component. [19] The method of claim 11, wherein a hardware reset of the component is performed by the data storage device and not by the host. [20] Data storage device comprising: a variety of memory dies means for determining that one or more memory dies of the plurality of memory dies are non-functional; means for informing a host communicating with the data storage device that an extended rebuild operation is being performed on the one or more storage dies such that a response to a host command may be delayed; and Means for performing the extended recovery operation on the one or more memory die(s).
Citation Information
Patent Citations
US-ANMELDUNGNR.63/449,770
US-ANMELDUNGNR.18/223,122