Data storage device and method for enhanced recovery by hardware reset of one of its discrete components

By configuring a hardware reset mechanism in the controller of the data storage device, the recovery problem of memory die is solved when the memory die is unresponsive, and the effect of reducing die back and improving performance and capacity is achieved.

CN120035807APending Publication Date: 2025-05-23SANDISK TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072743.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-18
Filing Date
2023-12-19
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When existing data storage devices are unresponsive to the memory die, it is difficult to effectively recover, resulting in read failures and performance losses.

Method used

By configuring a hardware reset mechanism in the controller of the data storage device, in response to the unresponsive subset of the memory die, the host is notified and performed to restore the function of the die.

Benefits of technology

This solution can reduce the number of die withdrawals, improve the performance and capacity of storage devices, and reduce the risk of performance losses and oversupply.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035807A_ABST
    Figure CN120035807A_ABST
Patent Text Reader

Abstract

A data storage device and method are provided for enhancing recovery by data storage device discrete component hardware reset. In one embodiment, the data storage device determines that a subset of a plurality of memory dies is unresponsive, transmits a request to a host to accept a longer latency associated with the subset of the plurality of memory dies, power cycles the subset of the plurality of memory dies, and transmits the request to the host to accept the longer latency associated with the subset of the plurality of memory dies. And then notifying the host that the latency associated with the dies has recovered to a normal latency or the subset of the plurality of memory dies is inactive (in the case of unsuccessful recovery). Other embodiments are possible, and each of the embodiments may be used alone or in combination.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of all contents of U.S. Non-Provisional Application No. 18 / 223,122, filed with the U.S. Patent and Trademark Office on July 18, 2023, entitled “Data Storage Device and Method for Enhanced Recovery Through a Hardware Reset of One of Its Discrete Components,” which claims priority to U.S. Provisional Application No. 63 / 449,770, filed on March 3, 2023, and which is incorporated herein by reference for all purposes. Background Art

[0003] As storage utilization in the enterprise and cloud grows, the need for data correction methods and high-reliability data storage devices will only grow. Some data storage device controllers use various data protection mechanisms to help ensure that read failure rates are low and that data returned to the host does not contain integrity errors, bit error rate (UBER), or mean time between failures (MTBF). BRIEF DESCRIPTION OF THE DRAWINGS

[0004] Figure 1A is a block diagram of a data storage device of an embodiment.

[0005] Figure 1B is a block diagram of a storage module of an example embodiment.

[0006] Figure 1C is a block diagram of a hierarchical storage system according to an example embodiment.

[0007] Figure 2A is an example according to the embodiment Figure 1A 00106] A block diagram of components of a controller of a data storage device illustrated in FIG.

[0008] Figure 2B is an example according to the embodiment Figure 1A 0036] A block diagram of components of a memory data storage device illustrated in FIG.

[0009] Figure 3 is a block diagram of a host and a data storage device of an embodiment.

[0010] Figure 4 is a flow chart of a method of an embodiment for die recovery.

[0011] Figure 5is a flow chart of a method for an embodiment of enhancing recovery operations.

[0012] Figure 6 is a flow chart of a method for an embodiment of enhancing recovery operations.

[0013] Figure 7 is a flow diagram of a method for an embodiment of an enhanced recovery operation based on various components. DETAILED DESCRIPTION

[0014] Overview

[0015] As an introduction, the following relates to a data storage device and method for enhancing recovery through a hardware reset of one embodiment of discrete components thereof. In one embodiment, a data storage device is provided, including: a plurality of memory dies; and a controller. The controller is configured to: in response to determining that a subset of the plurality of memory dies is unresponsive, notify a host that the data storage device will perform a hardware reset on the subset of the plurality of memory dies; in response to receiving an acknowledgement from the host, perform a hardware reset on the subset of the plurality of memory dies; and notify the host after the hardware reset has been performed on the subset of the plurality of memory dies.

[0016] In some embodiments, a hardware reset is initiated by the data storage device rather than the host.

[0017] In some embodiments, a hardware reset is performed on a subset of the plurality of memory dies without clearing host input-output queues for the subset of the plurality of memory dies.

[0018] In some embodiments, the controller is further configured to receive, from the host, a memory access command redirected from a subset of the plurality of memory dies to another memory die in the plurality of memory dies.

[0019] In some embodiments, the controller is further configured to perform a hardware reset on a subset of the plurality of memory dies by: transmitting a command to all of the plurality of memory dies to ignore the hardware reset command, wherein the subset of the plurality of memory dies does not receive the command to ignore the hardware reset command because the subset of the plurality of memory dies is unresponsive; and transmitting the hardware reset command on a communication channel shared by the plurality of memory dies.

[0020] In some embodiments, the controller is further configured to: determine whether a subset of the multiple memory dies is found to be unresponsive more than a threshold number of times within a time period; in response to determining that the subset of the multiple memory dies is not found to be unresponsive more than a threshold number of times within the time period, perform a hardware reset on the subset of the multiple memory dies; and in response to determining that the subset of the multiple memory dies is found to be unresponsive more than a threshold number of times within the time period, retire the subset of the multiple memory dies.

[0021] In some embodiments, the controller is further configured to, in response to determining that the host has enabled enhanced recovery operations, notify the host that the data storage device is to perform a hardware reset.

[0022] In some embodiments, a subset of the plurality of memory dies includes a single memory die.

[0023] In some embodiments, a subset of the plurality of memory dies includes more than one memory die.

[0024] In some embodiments, the plurality of memory dies includes three-dimensional memory dies.

[0025] In another embodiment, a method is provided for execution in a data storage device in communication with a host. The method includes: determining that a component in the data storage device needs to be hardware reset; notifying the host that the data storage device will perform a hardware reset of the component; performing the hardware reset of the component; and notifying the host after the hardware reset of the component has been performed.

[0026] In some embodiments, the component includes a memory die.

[0027] In some embodiments, the component comprises a portion of a controller of a data storage device.

[0028] In some implementations, the components include one or more of: dynamic random access memory (DRAM), electrically erasable programmable read-only memory (EEPROM), components in an application specific integrated circuit (ASIC), capacitors, and a bus controller.

[0029] In some embodiments, the data storage device determines that a component requires a hardware reset in response to determining that the component is unresponsive.

[0030] In some embodiments, the data storage device determines that a component requires a hardware reset in response to determining that the component exhibits erratic power usage.

[0031] In some embodiments, the data storage device determines that a component requires a hardware reset in response to determining that the component exhibits certain read and / or write latency characteristics.

[0032] In some embodiments, the method further includes receiving an instruction from a host to enable a function in the data storage device to perform a hardware reset of the component.

[0033] In some embodiments, a component is hardware reset by the data storage device rather than the host.

[0034] In another embodiment, a data storage device is provided, comprising: a plurality of memory dies; a device for determining that one or more of the plurality of memory dies are not functioning; a device for notifying a host communicating with the data storage device that an enhanced recovery operation will be performed on the one or more memory dies, thereby possibly delaying a response to a host command; and a device for performing an enhanced recovery operation on the one or more memory dies.

[0035] Other embodiments are possible, and each of the embodiments may be used alone or in combination.Therefore, various embodiments will now be described with reference to the accompanying drawings.

[0036] Implementation

[0037] The following embodiments relate to a data storage device (DSD). As used herein, a "data storage device" refers to a device that stores data. Examples of DSDs include (but are not limited to) hard disk drives (HDDs), solid-state drives (SSDs), tape drives, hybrid drives, and the like. Details of an example DSD are provided below.

[0038] Figures 1A to 1C A data storage device suitable for implementing aspects of these embodiments is shown in FIG. Figure 1A is a block diagram illustrating a data storage device 100 according to an embodiment of the subject matter described herein. Figure 1A , the data storage device 100 includes a controller 102 and a non-volatile memory that may be composed of one or more non-volatile memory die 104. As used herein, the term die refers to a collection of non-volatile memory cells formed on a single semiconductor substrate and associated circuitry for managing the physical operation of those non-volatile memory cells. The controller 102 interfaces with the host system and sends command sequences for read, program, and erase operations to the non-volatile memory die 104.

[0039] The controller 102 (which may be a non-volatile memory controller (e.g., a flash memory, a resistive random access memory (ReRAM), a phase change memory (PCM), or a magnetoresistive random access memory (MRAM) controller)) may take the form of (e.g.) a processing circuit, a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. The controller 102 may be configured with hardware and / or firmware to perform the various functions described below and shown in the flow chart. Moreover, some components shown as being internal to the controller may also be stored external to the controller, and other components may be used. In addition, the phrase "operably communicating with..." may mean communicating directly with... or communicating indirectly (wired or wirelessly) with... through one or more components, which may or may not be shown or described herein.

[0040] As used herein, a nonvolatile memory controller is a device that manages data stored on a nonvolatile memory and communicates with a host (such as a computer or electronic device). A nonvolatile memory controller may have various functionalities in addition to the specific functionality described herein. For example, a nonvolatile memory controller may format a nonvolatile memory to ensure that the memory operates correctly, map out bad nonvolatile memory cells, and allocate spare cells to replace future failed cells. A portion of the spare cell may be used to maintain firmware to operate the nonvolatile memory controller and implement other features. In operation, when the host needs to read data from the nonvolatile memory or write data to the nonvolatile memory, it may communicate with the nonvolatile memory controller. If the host provides a logical address where the data will be read / written, the nonvolatile memory controller may convert the logical address received from the host into a physical address in the nonvolatile memory. (Alternatively, the host may provide a physical address). The non-volatile memory controller may also perform various memory management functions such as (but not limited to) wear leveling (distributing writes to avoid wearing out specific blocks of memory that would otherwise be written repeatedly) and garbage collection (after a block is full, only valid data pages are moved to a new block so the full block can be erased and reused).

[0041] The non-volatile memory die 104 may include any suitable non-volatile storage medium, including resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), phase change memory (PCM), NAND flash memory cells, and / or NOR flash memory cells. The memory cells may take the form of solid-state (e.g., flash) memory cells and may be one-time programmable, several-time programmable, or multiple-time programmable. The memory cells may also be single-level cells (SLC), multi-level cells (MLC) (e.g., dual-level cells, triple-level cells (TLC), quad-level cells (QLC), etc.), or use other memory cell level technologies now known or later developed. Moreover, the memory cells may be manufactured in two or three dimensions.

[0042] The interface between the controller 102 and the non-volatile memory die 104 may be any suitable flash memory interface, such as switching mode 200, 400, or 800. In one embodiment, the data storage device 100 may be a card-based system, such as a secure digital (SD) or micro secure digital (micro-SD) card. In an alternative embodiment, the data storage device 100 may be part of an embedded data storage device.

[0043] Despite Figure 1A In the example illustrated in , data storage device 100 (sometimes referred to herein as a memory module) includes a single channel between controller 102 and non-volatile memory die 104, but the subject matter described herein is not limited to having a single memory channel. For example, in some architectures (such as Figure 1B and Figure 1C In the architecture shown in FIG. 1 , there may be two, four, eight, or more memory channels between the controller and the memory devices, depending on the controller capabilities. In any of the embodiments described herein, there may be more than a single channel between the controller and the memory die, even though a single channel is shown in the figure.

[0044] Figure 1B A storage module 200 including a plurality of non-volatile data storage devices 100 is illustrated. Thus, the storage module 200 may include a storage controller 202 that interfaces with a host and with a data storage device 204 including a plurality of data storage devices 100. The interface between the storage controller 202 and the data storage device 100 may be a bus interface, such as a Serial Advanced Technology Attachment (SATA), a Peripheral Component Interconnect Express (PCIe) interface, or a Double Data Rate (DDR) interface. In one embodiment, the storage module 200 may be a solid state drive (SSD) or a non-volatile dual in-line memory module (NVDIMM), such as found in a server PC or a portable computing device such as a laptop computer and a tablet computer.

[0045] Figure 1C 2 is a block diagram illustrating a hierarchical storage system. Hierarchical storage system 250 includes multiple storage controllers 202, each of which controls a corresponding data storage device 204. A host system 252 can access the memory within storage system 250 via a bus interface. In one embodiment, the bus interface can be a high-speed non-volatile memory (NVMe) or Ethernet Fibre Channel (FCoE) interface. In one embodiment, Figure 1C The system illustrated in may be a rack-mountable mass storage system accessible by multiple host computers, such as found in a data center or other location where mass storage is required.

[0046] Figure 2A 1 is a block diagram illustrating the components of the controller 102 in more detail. The controller 102 includes a front-end module 108 that interfaces with a host, a back-end module 110 that interfaces with one or more non-volatile memory dies 104, and various other modules that perform functions that will now be described in detail. For example, a module can take the form of a packaged functional hardware unit designed for use with other components, a portion of a program code (e.g., software or firmware) that can be executed by a (micro)processor or processing circuit that generally performs a specific one of the related functions, or a self-contained hardware or software component that interfaces with a larger system. In addition, a "means" for performing a certain function can be implemented with at least any structure for a controller described herein, and can be pure hardware or a combination of hardware and computer-readable program code.

[0047] Referring again to the modules of controller 102, buffer manager / bus controller 114 manages the buffers in random access memory (RAM) 116 and controls the internal bus arbitration of controller 102. Read only memory (ROM) 118 stores system boot code. Figure 2A 102, but in other embodiments, one or both of RAM 116 and ROM 118 may be located within the controller. In other embodiments, portions of RAM and ROM may be located within the controller 102 and outside the controller.

[0048] The front-end module 108 includes a host interface 120 and a physical layer interface (PHY) 122 that provide an electrical interface with a host or a next-level storage controller. The selection of the type of host interface 120 may depend on the type of memory used. Examples of host interface 120 include, but are not limited to, SATA, SATA Express, Serial Attached Small Computer System Interface (SAS), Fibre Channel, Universal Serial Bus (USB), PCIe, and NVMe. The host interface 120 generally facilitates the transmission of data, control signals, and timing signals.

[0049] The back-end module 110 includes an error correction code (ECC) engine 124 that encodes data bytes received from the host and decodes and corrects data bytes read from the non-volatile memory. A command sequencer 126 generates a command sequence to be sent to the non-volatile memory die 104, such as a program and erase command sequence. A RAID (Redundant Array of Independent Drives) module 128 manages the generation of RAID parity and the recovery of failed data. RAID parity can be used as an additional level of integrity protection for data being written to the memory device 104. In some cases, the RAID module 128 can be part of the ECC engine 124. The memory interface 130 provides command sequences to the non-volatile memory die 104 and receives status information from the non-volatile memory die 104. In one embodiment, the memory interface 130 can be a double data rate (DDR) interface, such as a switching mode 200, 400 or 800 interface. The flash control layer 132 controls the overall operation of the back-end module 110.

[0050] The data storage device 100 also includes other discrete components 140, such as external electrical interfaces, external RAM, resistors, capacitors, or other components that can interface with the controller 102. In alternative embodiments, one or more of the physical layer interface 122, the RAID module 128, the media management layer 138, and the buffer management / bus controller 114 are optional components that are not necessary in the controller 102.

[0051] Figure 2B 1 is a block diagram illustrating components of the nonvolatile memory die 104 in greater detail. The nonvolatile memory die 104 includes peripheral circuitry 141 and a nonvolatile memory array 142. The nonvolatile memory array 142 includes nonvolatile memory cells for storing data. The nonvolatile memory cells may be any suitable nonvolatile memory cells, including ReRAM, MRAM, PCM, NAND flash memory cells, and / or NOR flash memory cells in a two-dimensional and / or three-dimensional configuration. The nonvolatile memory die 104 also includes a data cache 156 that caches data. The peripheral circuitry 141 includes a state machine 152 that provides state information to the controller 102.

[0052] Back again Figure 2A, the flash control layer 132 (referred to herein as the flash translation layer (FTL), or more generally as the "media management layer" because the memory may not be flash) handles flash errors and interfaces with the host. In particular, the FTL (which may be an algorithm in firmware) is responsible for the internals of memory management and converts writes from the host into writes to the memory 104. The FTL may be needed because the memory 104 may have limited endurance, may only be written in multiple pages, and / or may not be written unless it is erased as a block. The FTL understands these potential limitations of the memory 104, which may not be visible to the host. Therefore, the FTL attempts to convert writes from the host into writes into the memory 104.

[0053] The FTL may include a logical to physical address (L2P) mapping (sometimes referred to herein as a table or data structure) and an allocated cache memory. In this way, the FTL converts a logical block address ("LBA") from the host into a physical address in memory 104. The FTL may include other features such as (but not limited to) power failure recovery (so that the FTL's data structures can be restored in the event of a sudden power failure) and wear leveling (so that wear across memory blocks is even to prevent certain blocks from being excessively worn, which would result in a greater likelihood of failure).

[0054] Turning again to the attached figure, Figure 3 300 and data storage device 100 of the embodiment. Host 300 can take any suitable form, including but not limited to a computer, a mobile phone, a tablet computer, a wearable device, a digital video recorder, a monitoring system, etc. The host 300 (here a computing device) in this embodiment includes a processor 330 and a memory 340. In one embodiment, the computer readable program code stored in the host memory 340 configures the host processor 330 to perform the actions described herein. Therefore, the actions performed by the host 300 are sometimes referred to as being performed by an application (computer readable program code) running on the host 300 in this article. For example, the host 300 can be configured to transfer data (e.g., initially stored in the host's memory 340) to the data storage device 100 for storage in the memory 104 of the data storage device.

[0055] As described above, the data storage device controller may use various data protection mechanisms to ensure low read failure rates and that the data returned to the host does not contain integrity errors. One such common mechanism in flash data storage devices is exclusive OR (XOR) protection for block-level or die-level recovery, which is used in high-capacity products that require low unrecoverable bit error rates (UBER) or mean time between failures (MTBF). These elements may include error detection and correction on storage / memory elements located in the data storage device controller and on the non-volatile memory die. Connection problems between the controller and the memory die may result in integrity errors that are monitored and detected in the controller.

[0056] When an error in memory is detected, a read failure often occurs when the data is deemed unrecoverable. In server-based flash storage devices, when the die is unresponsive, it may be removed from the system and an attempt to recover the data may be made through an XOR mechanism, which may incur a performance penalty during the recovery process. This process is Figure 4 is shown in the flowchart 400. Figure 4 As shown, it is detected that the die is unresponsive (action 410). Then, a die removal process is initiated, in which the data is recovered through an exclusive OR (XOR) operation and written to another die (action 420). The end result is that the die is removed without data being written or read to it (action 430). As described above, during the die removal process, the host may experience a performance abrupt change due to a large number of read recovery operations. In addition, depending on the number of dies removed, over-provisioning (OP) and therefore performance degradation may also be reduced.

[0057] This recovery scheme may be applied only at very high capacity points where the loss of a die can be accommodated. The reason for retiring a die when it is unresponsive is to attempt to recover dynamically with minimal impact to the quality of service (QoS) of ongoing input-output operations (IO). Moreover, in some cases, there is no way to restore communications with the die. However, in many cases, the cause of the die being unresponsive is a transient corruption in its internal state machine, which can be cleared with a power cycle (i.e., cutting and turning on power to the memory), and thereafter the component can function correctly without the additional risk of data loss.

[0058] The following embodiments can be used to enhance the recovery of memory elements with internal reset / power cycles of inoperative memory die through a protocol agreed upon with the host. This can result in a greatly reduced likelihood of a negative impact from die removal. For example, performing a hardware reset of the memory die (e.g., power cycling or performing a hard reset) can clear information stored in the latches / cache of the memory die, state machine information, interface settings / information, etc., which may be corrupted and cause memory die failure. This can recover data storage device components (e.g., NAND die) that cannot be recovered by methods such as XOR or software reset. This can also help reduce the number of retired dies and help improve performance, over-provisioning, and capacity.

[0059] In one embodiment, when a system element (e.g., a memory die) is unresponsive, the controller 102 in the data storage device 100 may attempt to reset it through a series of steps after establishing an agreement with the host 300. When the memory die 104 appears unresponsive, through a handshake with the host 300, the controller 102 of the data storage device 100 may indicate its recovery attempt (e.g., a warning of a short IO pause) but avoid retirement of the target die 104. This does not require the host 300 to run a power cycle or reinitialize the NVMe configuration. This also does not require the host 300 to clear the queue IO of the drive. The host 300 can approve the use of enhanced recovery for each of the supported components (it can select the components for which it will support enhanced recovery).

[0060] This implementation plan Figure 5 500. Figure 5 As shown, after detecting no response (act 510), the controller 102 of the data storage device 100 indicates to the host 300 that enhanced recovery is required (act 520). In response, the host 300 stops sending operations to the memory die or accepts a greater delay and notifies the data storage device 100 that it can begin the enhanced recovery process (act 530). The controller 102 of the data storage device 100 then performs enhanced recovery by performing a hardware reset of the memory die 104 and notifies the host 300 when the execution is completed (act 540). The host 300 then resumes operation of the subject data storage device 100 (act 550).

[0061] As used herein, "hardware reset" may refer to a power cycle operation or hard reset that restores a component to the state it was in when it left the factory. With a hardware reset, settings, applications, and user data may be removed. In contrast, a "firmware (or soft) reset" may refer to restarting a component to clear data from volatile memory and restart applications without completely shutting down the component.

[0062] Thus, in this embodiment, when the controller 102 first notifies the host 300 that enhanced recovery needs to be performed and a certain memory die (or several dies) will be unavailable (i.e., the host 300 should avoid transmitting memory access commands to the unresponsive memory die or accepting a large delay), a handshake between the host 300 and the controller 102 occurs. In response, the host 300 may stop all memory access (IO) operations (e.g., read and / or write commands) to the memory die that is undergoing a reset driven by a data storage device. For example, the host 300 may be prepared for IO suspension (e.g., by waiting for the completion of pending IO and redirecting the next IO to other memory dies) and respond when it is ready for the controller 102 to perform an enhanced recovery operation to reset the memory die. The host 300 does not need to perform a power cycle or reinitialize the NVMe configuration because the controller 102 can power cycle the memory die itself. The host 300 also does not need to clear the queue. Once enhanced recovery is completed, the host 300 can continue as usual.

[0063] Also, in some memory architectures (such as NAND), the reset pin (or other type of communication channel) is shared between all memory dies (e.g., when assuming a large number of dies in a device, a per-die pin or address-based system may not be feasible). Assuming that only a subset of the memory dies (e.g., one) may need to be reset, a command may be transmitted to the other memory dies to ignore the reset pin. In this way, the controller 102 can direct the reset to only specific memory dies. Inoperative dies cannot receive any commands, so the "ignore" command will only be received by the operational memory dies.

[0064] In another embodiment, there may be some heuristic that if the die appears "unresponsive" multiple times within a certain period of time, the previous retirement protocol is triggered. Figure 6 is shown in the flowchart 600. Figure 6 As shown, after having detected that the die is unresponsive (act 610), the controller 102 determines whether the die has shown no response within a period of time (e.g., within the last X minutes) (act 620). (Act 620 may be modified with another heuristic that takes into account irregular behavior of the subject die other than "no response", such as irregular power, read / write latency characteristics, or any other characteristics that indicate that a component in the data storage device 100 is not functioning or is experiencing a failure). If the die does not show no response within the period of time, the controller 102 performs an enhanced recovery operation using the host handshake (e.g., such as Figure 5 ) (act 630). However, if the die has shown no response within the time period, controller 102 retires the die (act 640).

[0065] In another embodiment, other elements within the storage controller 102 may trigger enhanced recovery operations. For example, portions of the dynamic random access memory ("DRAM"), electrically erasable programmable read-only memory ("EEPROM"), or the ASIC itself may be turned off and on (or reset) to recover from transient errors associated with the memory or interfaces. Additionally, enhanced recovery may be applied to non-memory die / memory elements such as (but not limited to) capacitors or a system management bus (SMBus) controller. Furthermore, enhanced recovery may be enabled or disabled individually by the host 300 for each of these components, potentially based on the expected IO pause duration reported by the data storage device 100. The host 300 may also direct the drive to apply recovery internally without additional notification to the host 300.

[0066] When the storage controller 102 identifies a problem with a recoverable component, it can check whether the host 300 has enabled enhanced recovery for the corresponding component. Based on the component, it can schedule the execution of recovery operations. For example, if the DRAM is inaccessible in the corresponding cycle, the controller 102 can still access the local storage device if there are some urgent host operations to be completed.

[0067] Figure 7 A flow chart 700 illustrating this embodiment is provided. Figure 7 As shown, the host 300 configures (enables / disables) enhanced recovery for different data storage device components (action 710). Component failure is detected using the components enabled for enhanced recovery (action 720). The controller 102 of the data storage device 100 then indicates to the host 300 the need for enhanced recovery of the corresponding component (action 730). The host 300 operates according to the failed component and notifies the data storage device 100 that it can start the enhanced recovery process (action 740). The controller 102 of the data storage device 100 then performs the enhanced recovery and notifies the host 300 when the execution is completed (action 750). The host 300 then resumes normal operation of the data storage device 100 (action 760).

[0068] Therefore, in this embodiment, the host 300 configures the components and potentially configures the corresponding allowed recovery duration, which may also determine whether the controller 102 initiates an enhanced recovery handshake based on the nature of the failure. Then, upon detecting a failure of a component approved by the host 300 for enhanced recovery, the controller 102 may initiate a handshake process with the host 300. The host 300's reaction operation may depend on Figure 7 The nature of the faulty component in action 740 in .

[0069] These embodiments have several advantages. For example, the embodiments can reduce the number of retired dies, improve performance, over-provisioning and capacity, and improve storage devices in high-end server products.

[0070] Finally, as described above, any suitable type of memory may be used. Semiconductor memory devices include volatile memory devices (such as dynamic random access memory ("DRAM") or static random access memory ("SRAM") devices), non-volatile memory devices (such as resistive random access memory ("ReRAM"), electrically erasable programmable read-only memory ("EEPROM"), flash memory (which may be considered a subset of EEPROM), ferroelectric random access memory ("FRAM"), and magnetoresistive random access memory ("MRAM"), as well as other semiconductor elements capable of storing information. Each type of memory device may have a different configuration. For example, a flash memory device may be configured in a NAND or NOR configuration.

[0071] The memory device may be formed of passive and / or active elements in any combination. As non-limiting examples, passive semiconductor memory elements include ReRAM device elements, which in some embodiments include resistivity switching storage elements such as antifuses, phase change materials, etc., and optionally include steering elements such as diodes, etc. As non-limiting examples, active semiconductor memory elements include EEPROM and flash memory device elements, which in some embodiments include elements including charge storage regions such as floating gates, conductive nanoparticles, or charge storage dielectric materials.

[0072] A plurality of memory elements may be configured so that they are connected in series or so that each element is individually accessible. As a non-limiting example, a flash memory device (NAND memory) in a NAND configuration typically contains memory elements connected in series. A NAND memory array may be configured so that the array is composed of a plurality of memory strings, wherein the string is composed of a plurality of memory elements that share a single bit line and are accessed as a group. Alternatively, the memory element may be configured so that each element is individually accessible, for example, a NOR memory array. NAND and NOR memory configurations are examples, and the memory element may be configured in other ways.

[0073] The semiconductor memory elements located in and / or on the substrate may be arranged in two dimensions or three dimensions, such as a two-dimensional memory structure or a three-dimensional memory structure.

[0074] In a two-dimensional memory structure, semiconductor memory elements are arranged in a single plane or a single memory device level. Typically, in a two-dimensional memory structure, the memory elements are arranged in a plane extending substantially parallel to the main surface of the substrate supporting the memory elements (e.g., in an xz-direction plane). The substrate may be a wafer on which or in which a memory element layer is formed, or the substrate may be a carrier substrate attached to the memory element after the memory element is formed. As a non-limiting example, the substrate may include a semiconductor, such as silicon.

[0075] The memory elements may be arranged in an ordered array (such as in multiple rows and / or columns) in a single memory device level. However, the memory elements may be arranged in an irregular or non-orthogonal configuration. The memory elements may each have two or more electrodes or contact lines, such as a bit line and a word line.

[0076] A three-dimensional memory array is arranged such that the memory elements occupy multiple planes or multiple memory device levels, thereby forming a three-dimensional structure (i.e., in the x, y and z directions, where the y direction is generally perpendicular to the major surface of the substrate and the x and z directions are generally parallel to the major surface of the substrate).

[0077] As a non-limiting example, a three-dimensional memory structure may be arranged vertically as a stack of multiple two-dimensional memory device levels. As another non-limiting example, a three-dimensional memory array may be arranged as a plurality of vertical columns (e.g., columns extending generally perpendicular to the major surface of the substrate (i.e., in the y-direction)), wherein each column has a plurality of memory elements in each column. The columns may be arranged in a two-dimensional configuration (e.g., in the xz plane), thereby producing a three-dimensional arrangement of memory elements having multiple vertically stacked elements on the memory plane. Other configurations of memory elements in three dimensions may also constitute a three-dimensional memory array.

[0078] As a non-limiting example, in a three-dimensional NAND memory array, memory elements may be coupled together to form NAND strings within a single horizontal (e.g., xz) memory device level. Alternatively, memory elements may be coupled together to form vertical NAND strings that span multiple horizontal memory device levels. Other three-dimensional configurations are contemplated, with some NAND strings containing memory elements in a single memory level and other strings containing memory elements that span multiple memory levels. Three-dimensional memory arrays may also be designed in NOR configurations and ReRAM configurations.

[0079] Typically, in a monolithic three-dimensional memory array, one or more memory device levels are formed over a single substrate. Optionally, a monolithic three-dimensional memory array may also have one or more memory layers at least partially within a single substrate. As a non-limiting example, the substrate may include a semiconductor, such as silicon. In a monolithic three-dimensional array, the layers of each memory device level that make up the array are typically formed on the layers of the underlying memory device level of the array. However, the layers of adjacent memory device levels of a monolithic three-dimensional memory array may share or have intervening layers between the memory device levels.

[0080] Likewise, two-dimensional arrays may be formed separately and then packaged together to form a non-monolithic memory device having multiple layers of memory. For example, a non-monolithic stacked memory may be constructed by forming memory levels on separate substrates and then stacking the memory levels on top of each other. The substrate may be thinned or removed from the memory device levels before stacking, but since the memory device levels are initially formed on separate substrates, the resulting memory array is not a monolithic three-dimensional memory array. In addition, multiple two-dimensional memory arrays or three-dimensional memory arrays (monolithic or non-monolithic) may be formed on separate chips and then packaged together to form a stacked chip memory device.

[0081] The operation of the memory element and the communication with the memory element generally require associated circuits. As a non-limiting example, the memory device may have circuits for controlling and driving the memory element to implement functions such as programming and reading. Such associated circuits may be on the same substrate as the memory element and / or on a separate substrate. For example, a controller for memory read-write operations may be located on a separate controller chip and / or on the same substrate as the memory element.

[0082] Those skilled in the art will recognize that the present invention is not limited to the two-dimensional and three-dimensional structures described, but encompasses all related memory structures within the spirit and scope of the invention as described herein and as understood by those skilled in the art.

[0083] The above specific embodiments are intended to be understood as illustrations of selected forms that the present invention may take, rather than limitations of the present invention. Only the following claims (including all equivalents) are intended to define the scope of the present invention as claimed. Finally, it should be noted that any aspect of any embodiment described herein may be used alone or in combination with each other.

Claims

1. A data storage device, the data storage device include: a plurality of memory dies; and a controller configured to communicate with the plurality of memory dies and further configured to: In response to determining that a subset of the plurality of memory dies are unresponsive, notifying a host that the data storage device will perform a hardware reset on the subset of the plurality of memory dies; in response to receiving an acknowledgement from the host, performing the hardware reset on the subset of the plurality of memory dies; as well as The host is notified after the hardware reset has been performed on the subset of the plurality of memory dies. 2 . The data storage device of claim 1 , wherein the hardware reset is initiated by the data storage device rather than the host.

3. The data storage device of claim 1, wherein the subset of the plurality of memory dies is hardware reset without clearing a host input-output queue for the subset of the plurality of memory dies. 4 . The data storage device of claim 1 , wherein the controller is further configured to receive from the host a memory access command redirected from the subset of the plurality of memory dies to another memory die of the plurality of memory dies.

5. The data storage device of claim 1 , wherein the controller is further configured to perform the hardware reset on the subset of the plurality of memory dies by: transmitting a command to all of the plurality of memory dies to ignore a hardware reset command, wherein the subset of the plurality of memory dies does not receive the command to ignore the hardware reset command because the subset of the plurality of memory dies is non-responsive; and The hardware reset command is transmitted over a communication channel shared by the plurality of memory dies.

6. The data storage device of claim 1 , wherein the controller is further configured to: determining whether the subset of the plurality of memory dies is found to be unresponsive more than a threshold number of times within a time period; in response to determining that the subset of the plurality of memory die has not been found to be unresponsive more than the threshold number of times within the time period, performing the hardware reset on the subset of the plurality of memory die; as well as In response to determining that the subset of the plurality of memory die was found to be unresponsive more than the threshold number of times within the time period, the subset of the plurality of memory die is retired. 7 . The data storage device of claim 1 , wherein the controller is further configured to, in response to determining that the host has enabled enhanced recovery operations, notify the host that the data storage device will perform the hardware reset.

8. The data storage device of claim 1, wherein the subset of the plurality of memory die comprises a single memory die.

9. The data storage device of claim 1, wherein the subset of the plurality of memory die comprises more than one memory die.

10. The data storage device of claim 1, wherein the plurality of memory dies comprises three-dimensional memory dies.

11. A method, the method comprising: include: In a data storage device that communicates with the host, perform the following operations: determining that a component in the data storage device requires a hardware reset; notifying the host that the data storage device will perform a hardware reset of the component; performing said hardware reset of said component; as well as The host is notified after the hardware reset of the component has been performed.

12. The method of claim 11, wherein the component comprises a memory die.

13. The method of claim 11, wherein the component comprises a portion of a controller of the data storage device.

14. The method of claim 11, wherein the components include one or more of: dynamic random access memory (DRAM), electrically erasable programmable read-only memory (EEPROM), components in an application specific integrated circuit (ASIC), capacitors, and a bus controller.

15. The method of claim 11, wherein the data storage device determines that the component requires a hardware reset in response to determining that the component is unresponsive.

16. The method of claim 11, wherein the data storage device determines that the component requires a hardware reset in response to determining that the component exhibits erratic power usage.

17. The method of claim 11, wherein the data storage device determines that the component requires a hardware reset in response to determining that the component exhibits certain read and / or write latency characteristics.

18. The method of claim 11, further comprising receiving an instruction from the host to enable a function in the data storage device to perform a hardware reset on the component.

19. The method of claim 11, wherein the component is hardware reset by the data storage device rather than the host.

20. A data storage device, the data storage device include: a plurality of memory dies; means for determining that one or more memory dies of the plurality of memory dies are non-functional; means for notifying a host in communication with the data storage device that an enhanced recovery operation will be performed on the one or more memory dies, thereby potentially delaying a response to a host command; as well as Means for performing the enhanced recovery operation on the one or more memory dies.