Host control mitigation in panic situations

The integration of panic early detection and control modules in data storage devices enables proactive mitigation of potential failures by offering detailed data and options to the host, enhancing system resilience.

JP2026047079APending Publication Date: 2026-03-13SANDISK TECHNOLOGIES LLC
View PDF 10 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Current data storage devices lack the capability to prepare in advance for future panic situations, leading to potential failures without adequate mitigation strategies.

Method used

Implementing a panic early detection module and control module in data storage devices to detect future panic conditions, providing detailed panic data and mitigation options to the host device, such as adjusting performance, increasing power, or redirecting commands to another storage device.

Benefits of technology

Enhances the robustness of data storage systems by allowing for proactive mitigation of panic conditions, reducing the impact of failures and maintaining system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026047079000001_ABST
    Figure 2026047079000001_ABST
Patent Text Reader

Abstract

This invention provides a data storage device that reduces data storage device failure by detecting panic situations and providing the host device with detailed panic data and mitigation options before panic conditions occur. [Solution] When a future panic situation is detected, the data storage device presents the host device with several mitigation options, such as adjusting read performance, increasing device power, performing evacuation and management operations, and / or redirecting the host command to another data storage device for command completion. When the other data storage device is in the same PCIe tree and reachable, the command is directed using PRP / SGL which again points to the same host device. The data storage devices may have transmit queues between them and include a panic early detection module and a panic control module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to data storage devices for early detection and mitigation of panic situations.

Background Art

[0002] Description of Related Art A device panic situation or panic situation is a situation in which, when a panic state (e.g., a failure) occurs, a data storage device (e.g., a solid state drive (SSD)) can notify a host device of mitigation steps to be taken. The panic state may be signaled using an asynchronous event or a controller error status register. When a panic state such as a failure occurs, a reset action may be performed by the host device. After the host device has performed mitigation steps corresponding to the identified panic state, the storage device is expected to provide diagnostic information. In some situations, the data storage device may also propose post (potential) reset actions to be taken.

[0003] When a data storage device detects a panic situation, the panic situation is reported to the host device through an interface, and the capabilities of the data storage device are provided to the host device during the panic situation. However, currently, there are no requirements or methods for the data storage device to prepare in advance for future panic situations before an actual panic situation occurs.

[0004] Therefore, there is a need for a technology for improved data storage devices for early detection and mitigation of panic situations.

Summary of the Invention

[0005] Failure of data storage devices can be mitigated by detecting panic conditions and providing the host device with detailed panic data and mitigation options before panic conditions occur. When a future panic condition is detected, several mitigation options may be presented to the host device, such as adjusting read performance, increasing device power, performing evacuation and management actions, and / or redirecting host commands to another data storage device for command completion. When other data storage devices are in the same PCIe tree and reachable, commands may be directed using PRP / SGL pointing again to the same host device. In some embodiments, data storage devices may have transmit queues among them. In some embodiments, data storage devices include a panic early detection module and a panic control module.

[0006] In one embodiment, the data storage device includes a memory device and a controller coupled to the memory device, the controller being configured to detect a future panic situation of the data storage device, analyze the future panic situation, and propose at least one mitigation option to a host device, wherein the proposal includes a decision timeout; execute a default mitigation option while waiting to receive a selected mitigation option from the host device; and receive a selected mitigation option from the host device, which has been selected from the proposed at least one mitigation option, and execute the selected mitigation option.

[0007] In another embodiment, the data storage device includes a memory device and a controller coupled to the memory device, the controller being configured to detect a future panic situation of the data storage device based on a panic indicator, analyze the future panic situation, propose at least one mitigation option to the host device, one of which includes redirecting a host command to another location, execute the mitigation option, and determine that the future panic situation has been resolved.

[0008] In another embodiment, the data storage device includes means for storing data and a controller coupled to the means for storing data, the controller is configured to detect a future panic situation of the data storage device, receive a host command queued in the send queue of the host device, and, for completion, redirect the host command to a second data storage device, the second data storage device being in the same PCIe tree as the data storage device, and interrupt the host with a completion entry into the host completion queue of the host device. [Brief explanation of the drawing]

[0009] A more detailed description of the Disclosure, which is concisely summarized above, may be obtained by reference to embodiments, some of which are shown in the accompanying drawings, so that the above-mentioned features of the Disclosure may be understood in detail. However, it should be noted that the accompanying drawings show only typical embodiments of the Disclosure and should not be considered to limit its scope, as the Disclosure may allow for other equally valid embodiments. [Figure 1] This is a schematic block diagram showing a storage system in which, according to a particular embodiment, a data storage device may function as a storage device for a host device. [Figure 2]This table shows various panic reset and recovery actions for data storage devices according to several embodiments. [Figure 3] This is a schematic block diagram showing a memory system with early panic detection and control according to several embodiments. [Figure 4] This table shows various potential error injection types for debugging panic conditions in data storage devices, according to several embodiments. [Figure 5] This flowchart shows a method for detecting and mitigating panic conditions in a data storage device according to several embodiments. [Figure 6] This flowchart shows a method for detecting and mitigating panic conditions in a data storage device according to several embodiments. [Figure 7A] This is a schematic block diagram illustrating a memory system for detecting and mitigating future panic situations, according to several embodiments. [Figure 7B] Figure 7A is a flowchart showing methods for detecting and mitigating panic situations in memory devices.

[0010] For ease of understanding, the same reference numerals are used to designate identical elements common to the drawings where possible. Elements disclosed in one embodiment are intended to be usefully utilized in other embodiments without specific description. [Modes for carrying out the invention]

[0011] The following refers to embodiments of the Disclosure. However, it should be understood that the Disclosure is not limited to any specific embodiment described. Instead, any combination of the following features and elements, whether related to a different embodiment or not, is intended to implement and practice the Disclosure. Furthermore, embodiments of the Disclosure may achieve advantages over other possible solutions and / or prior art, but whether a particular advantage is achieved by a given embodiment does not limit the Disclosure. Accordingly, the following aspects, features, embodiments, and advantages are merely illustrative and shall not be considered elements or limitations of the appended claims unless expressly enumerated in the claims. Similarly, references to “the Disclosure” shall not be construed as a generalization of the subject matter of any invention disclosed herein and shall not be considered elements or limitations of the appended claims unless expressly enumerated in the claims.

[0012] Failure of data storage devices can be mitigated by detecting panic conditions and providing the host device with detailed panic data and mitigation options before panic conditions occur. When a future panic condition is detected, several mitigation options may be presented to the host device, such as adjusting read performance, increasing device power, performing evacuation and management actions, and / or redirecting host commands to another data storage device for command completion. When other data storage devices are in the same PCIe tree and reachable, commands may be directed using PRP / SGL pointing again to the same host device. In some embodiments, data storage devices may have transmit queues among them. In some embodiments, data storage devices include a panic early detection module and a panic control module.

[0013] Figure 1 is a schematic block diagram showing a storage system 100 having a data storage device 106 which may function as a storage device for a host device 104 according to a particular embodiment. For example, the host device 104 may store and retrieve data using non-volatile memory (NVM) 110 contained in the data storage device 106. The host device 104 includes host dynamic random access memory (DRAM) 138. In some examples, the storage system 100 may include multiple storage devices, such as the data storage device 106, which may operate as a storage array. For example, the storage system 100 may include multiple data storage devices 106 configured as a redundant array of inexpensive / independent disks (RAID) which collectively function as a high-capacity storage device for the host device 104.

[0014] The host device 104 may store data in and / or retrieve data from one or more storage devices, such as the data storage device 106. As shown in Figure 1, the host device 104 may communicate with the data storage device 106 via the interface 114. The host device 104 may include any of a wide range of devices, including a computer server, a network-attached storage (NAS) unit, a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone, a so-called "smart" pad, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, or other devices capable of sending or receiving data from a data storage device.

[0015] The host DRAM 138 may optionally include a host memory buffer (HMB) 150. The HMB 150 is a portion of the host DRAM 138 allocated to the data storage device 106 for exclusive use by the controller 108. For example, the controller 108 may store mapping data, buffered commands, logical-to-physical (L2P) tables, metadata, etc., in the HMB 150. In other words, the HMB 150 may be used by the controller 108 to store data that would normally be stored in the controller 108's internal memory, such as volatile memory 112, buffers 116, or static random access memory (SRAM). In an example where the data storage device 106 does not include DRAM (i.e., the optional DRAM 118), the controller 108 may use the HMB 150 as the DRAM for the data storage device 106.

[0016] The data storage device 106 includes a controller 108, an NVM 110, a power supply 111, volatile memory 112, an interface 114, a write buffer 116, and an optional DRAM 118. In some examples, the data storage device 106 may include additional components not shown in Figure 1 for clarity. For example, the data storage device 106 may include a printed circuit board (PCB) on which the components of the data storage device 106 are mechanically mounted and which includes conductive traces that electrically interconnect the components such as the data storage device 106. In some examples, the physical dimensions and connector configuration of the data storage device 106 may conform to one or more standard form factors. Some exemplary standard form factors include, but are not limited to, 3.5-inch data storage devices (e.g., HDD or SSD), 2.5-inch data storage devices, 1.8-inch data storage devices, peripheral component interconnect (PCI), PCI-extended (PCI-X), and PCI Express (PCIe) (e.g., PCIe x1, x4, x8, x16, PCIe MiniCard, MiniPCI, etc.). In some examples, the data storage device 106 may be directly coupled to the motherboard of the host device 104 (e.g., directly soldered or plugged into a connector).

[0017] Interface 114 may include either or both a data bus for exchanging data with the host device 104 and a control bus for exchanging commands with the host device 104. Interface 114 may operate according to any preferred protocol. For example, interface 114 may operate according to the non-volatile memory express (NVMe) protocol, etc. Interface 114 (e.g., the data bus, the control bus, or both) is electrically connected to the controller 108 to provide an electrical connection between the host device 104 and the controller 108, enabling data to be exchanged between the host device 104 and the controller 108. In some examples, the electrical connection of interface 114 may also allow a data storage device 106 to receive power from the host device 104. For example, as shown in Figure 1, a power supply 111 may receive power from the host device 104 via interface 114.

[0018] The NVM110 may include multiple memory devices or memory units. The NVM110 may be configured to store and / or retrieve data. For example, a memory unit of the NVM110 may receive data and messages from controller 108 instructing the memory unit to store the data. Similarly, a memory unit may receive messages from controller 108 instructing the memory unit to retrieve data. In some examples, each of the memory units may be called a die. In some examples, the NVM110 may include multiple dies (i.e., multiple memory units). In some examples, each memory unit may be configured to store relatively large amounts of data (e.g., 128MB, 256MB, 512MB, 1GB, 2GB, 4GB, 8GB, 16GB, 32GB, 64GB, 128GB, 256GB, 512GB, 1TB, etc.).

[0019] In some examples, each memory unit may include any type of non-volatile memory device, such as flash memory devices, phase-change memory (PCM) devices, resistive random-access memory (ReRAM) devices, magneto-resistive random-access memory (MRAM) devices, ferroelectric random-access memory (F-RAM), holographic memory devices, and any other type of non-volatile memory device.

[0020] The NVM 110 may include a plurality of flash memory devices or memory units. The NVM flash memory device may include a NAND or NOR-based flash memory device, and may store data based on the charge contained in the floating gate of the transistor of each flash memory cell. In the NVM flash memory device, the flash memory device may be divided into a plurality of dies, each die of the plurality of dies includes a plurality of physical blocks or logical blocks, and the plurality of physical blocks or logical blocks may be further divided into a plurality of pages. Each block of the plurality of blocks within a specific memory device may include a plurality of NVM cells. The rows of NVM cells may be electrically connected using word lines to define one of the plurality of pages. Each cell in each of the plurality of pages may be electrically connected to a respective bit line. Further, the NVM flash memory device may be a 2D or 3D device, and may be a single level cell (SLC), multi-level cell (MLC), triple level cell (TLC), or quad level cell (QLC). The controller 108 can write data to the NVM flash memory device at the page level, read data from the NVM flash memory device, and erase data from the NVM flash memory device at the block level.

[0021] The power supply 111 may supply power to one or more components of the data storage device 106. When operating in standard mode, the power supply 111 may supply power to one or more components using power provided by an external device, such as a host device 104. For example, the power supply 111 may supply power to one or more components using power received from the host device 104 via interface 114. In some examples, the power supply 111 may include one or more power storage components configured to supply power to one or more components when operating in shutdown mode, such as when power is no longer received from an external device. In this way, the power supply 111 may function as an onboard backup power supply. Some examples of one or more power storage components include, but are not limited to, capacitors, supercapacitors, and batteries. In some examples, the amount of power that can be stored by one or more power storage components may be a function of the cost and / or size (e.g., area / volume) of the one or more power storage components. In other words, as the amount of power stored by one or more power storage components increases, the cost and / or size of the one or more power storage components also increases.

[0022] The volatile memory 112 may be used by the controller 108 to store information. The volatile memory 112 may include one or more volatile memory devices. In some examples, the controller 108 can use the volatile memory 112 as a cache. For example, the controller 108 may store cached information in the volatile memory 112 until the cached information is written to the NVM 110. As shown in Figure 1, the volatile memory 112 may consume power received from the power supply 111. Examples of volatile memory 112 include, but are not limited to, random-access memory (RAM), dynamic random-access memory (DRAM), static RAM (SRAM), and synchronous dynamic RAM (SDRAM (e.g., DDR1, DDR2, DDR3, DDR3L, LPDDR3, DDR4, LPDDR4, etc.)). Similarly, an optional DRAM 118 may be used to store mapping data, buffered commands, logical-to-physical (L2P) tables, metadata, cached data, etc. In some examples, the data storage device 106 does not include an optional DRAM 118, and is therefore DRAM-less. In other examples, the data storage device 106 includes an optional DRAM 118.

[0023] Controller 108 may manage one or more operations of data storage device 106. For example, controller 108 may manage reading data from NVM 110 and / or writing data to NVM 110. In some embodiments, when data storage device 106 receives a write command from host device 104, controller 108 may initiate a data storage command to store data in NVM 110 and monitor the progress of the data storage command. Controller 108 may determine at least one operating characteristic of storage system 100 and store the at least one operating characteristic in NVM 110. In some embodiments, when data storage device 106 receives a write command from host device 104, controller 108 temporarily stores the data associated with the write command in internal memory or write buffer 116 before transmitting the data to NVM 110. Controller 108 may include a circuit or processor configured to execute a program for operating data storage device 106.

[0024] Controller 108 may include an optional second volatile memory 120. The optional second volatile memory 120 may be similar to volatile memory 112. For example, the optional second volatile memory 120 may be SRAM. Controller 108 may allocate a portion of the optional second volatile memory as a controller memory buffer (CMB) 122 to host device 104. CMB 122 may be directly accessed by host device 104. For example, instead of maintaining one or more transmit queues within host device 104, host device 104 may use CMB 122 to store one or more transmit queues that are normally maintained within host device 104. In other words, host device 104 may generate a command and store the generated command with or without associated data in CMB 122, and controller 108 may access CMB

[0025] Figure 2 is Table 200, which shows various panic reset and recovery actions for a data storage device in several embodiments. Table 200 is taken from the publicly available Open Compute Project (OCP) Datacenter Specification, named Datacenter NVMe® SSD Specification (Version 2.0). The data storage device may use bit fields to indicate potential reset actions that may need to be performed during or before a panic situation to prevent it. The data storage device may also use bit fields to indicate appropriate device recovery actions to be taken to deal with a panic situation (e.g., device panic state or panic mode). As described below, the probability of device failure may be reduced by providing the host device with additional mitigation options and detailed panic data before a device panic state occurs (if possible), and by providing the host device with a set of recovery / mitigation options for dealing with a device panic state. The set of mitigation options may include modifying a subset of different device capabilities, such as reducing performance, increasing power drawn, removing the option to read from a subset of dies, or redirecting host commands to another available storage device. In some embodiments, such a storage system analyzes the current situation based on inputs regarding the health of the storage device and external conditions, and outputs indications of potential panic conditions to the host using log pages or other suitable means for signaling.

[0026] In some embodiments, when a host device is provided with a set of recovery / mitigation options for handling device panic conditions, the data storage device may have reduced device capabilities for the duration of the device panic condition. Note that while the error recovery log is designed for data center storage devices (e.g., enterprise storage devices), the log may also be adapted for client SSDs. The disclosed embodiments are applicable to both types of SSDs, data center storage devices and client SSDs, although certain characteristics may vary. For example, capacitor or DRAM failures may not be applicable to client SSDs, while HMB failures may occur in clients but not in data center devices.

[0027] Figure 3 is a schematic block diagram showing a memory system 300 with early panic detection and control according to several embodiments. The memory system 300 includes a host device 302, a data storage controller 304, and an NVM die 310. The host device 302 may be the host 104 in Figure 1. The data storage controller 304 may be the controller 108 in Figure 1. The NVM die 310 may be implemented on the NVM 110 in Figure 1. The data storage controller 304 includes a panic early detection module (PDM) 306 and a panic control module (PCM) 308.

[0028] PDM306 is configured to detect panic situations. The goal is to detect future panic situations as early as possible. As a result, the system will detect future panic situations, but a certain false alarm rate will be assumed by the system if the panic situation is avoided. PCM308 is configured to provide the host device 302 with several mitigation options to modify the memory system behavior when PDM306 detects a future device panic condition, depending on the panic ID (e.g., error type injection in Figure 4).

[0029] Figure 4 is Table 400, which shows various potential error injection types for debugging panic conditions in data storage devices according to several embodiments. Table 400 is taken from the OCP Data Center Specification. Possible causes of a device panic condition can range from firmware failures to NVM failures. Panic IDs and associated causes vary on a case-by-case basis, and examples of various potential error injection types for debugging are shown in Table 400. Therefore, once a data storage device (e.g., data storage device 304 in Figure 3) detects a panic condition and determines a panic ID associated with the cause, the data storage device may map the panic condition, the failure state of the panic condition, or the panic ID to the corresponding error injection type in Table 400 for debugging.

[0030] Device panic situations, such as hardware malfunctions, may involve a wide range of characteristics. Detecting these characteristics by a data storage device may indicate the presence of a potential panic situation, but not all or just one of the characteristics necessarily triggers a panic situation or assertion. For example, when an error correction code (ECC) engine malfunctions, the malfunction may be detected early via a PDM (e.g., PDM306 in Figure 3) by noticing a decrease in performance or an increase in power consumption. In this situation, there may be choices presented to the host device. For example, without triggering a failure, the host device may choose to maintain reduced read performance while using the same amount of power, or to use more power while expecting the same performance.

[0031] In another example, in the case of NAND failure, one of the dies may not be written to. This can be detected early through monitoring if the number of cycles required for writing is unusually high. In these situations, the system is likely to back up all data to another die. However, this presents a trade-off for the host device. The host device may experience a decrease in write performance, but can still read from the storage device, albeit with some reduced reliability, until the defective die eventually becomes unreadable. Alternatively, the host device may also provide a time slot to the data storage device to perform management operations, such as backing up and management operations, as well as attempts to revive the die for write commands. After this duration, the host device may eliminate the issue of reduced reliability, leaving only the decrease in write performance if the die failed to recover successfully.

[0032] In yet another example, a portion of the DRAM may be suspected of being corrupted. As a result, the DRAM may use some dedicated ECCs, and as soon as the problem is recognized, the storage device may indicate an early panic condition flag (i.e., before the corruption is actually witnessed). Once the size of the corrupted DRAM is determined, the controller (e.g., controller 108 in Figure 1) may propose reducing performance or exported capacity, so that the controller has less data to "control" using the remaining DRAM. The controller may also propose a disable feature that relies on the DRAM or, if available, uses most of the host device's DRAM (e.g., the host device's HMB, such as HMB 150 in Figure 1). Alternatively, the controller may also decide to increase the ECC bits in cooperation with the DRAM to improve integrity. In some embodiments, the selection by the data storage device should be communicated to the controller via the same interface within a timeframe; otherwise, the data storage device selects a default option to mitigate the detected early panic condition.

[0033] Figure 5 is a flowchart illustrating a method 500 for detecting and mitigating panic situations in a data storage device, according to several embodiments. The method 500 begins in operation 502, when a PDM (e.g., PDM 306 in Figure 3) of a controller (e.g., controller 108 in Figure 108) monitors and detects potential panic situations before they occur. In operation 504, a PCM (e.g., PCM 308 in Figure 3) analyzes the panic situation and proposes several mitigation options for the panic situation, including the mitigation options described above. In operation 506, the early panic situation (e.g., the cause of the panic situation and other information regarding the panic situation) and the mitigation options are sent to the host device along with a timeout for the host device to make a decision. In operation 508, the data storage device determines whether the host device has made a decision on a mitigation option within the given timeout. If the host device has made a decision on a mitigation option within the given timeout by notifying the controller of the selected mitigation, in operation 510, the controller performs the correction specified by the selected mitigation option. If the mitigation option sent to the host device times out and the host device has not selected a mitigation option, in operation 512, the controller performs the correction specified by the default mitigation option.

[0034] Figure 6 is a flowchart illustrating a method 600 for detecting and mitigating panic conditions in a data storage device, according to several embodiments. In some embodiments, a PDM (e.g., PDM 306 in Figure 3) detects a potential panic condition before it occurs, and a PCM (e.g., PCM 308 in Figure 3) analyzes the available appropriate mitigation options for the detected panic condition. A controller (e.g., controller 108 in Figure 108) may then immediately switch from several appropriate mitigation options to a default mitigation option. In parallel, the appropriate mitigation options are also transmitted and posted to the host device, which can override the implemented default mitigation option by selecting a mitigation option. As a result, the data storage device is not exposed to further failures while the host selects a mitigation option, thereby promoting device health. In some embodiments, the panic condition may be temporary, and the controller attempts to resolve the panic condition. In these situations, the controller may use an interface (e.g., interface 114 in Figure 1) to remove panic indicators (which may require a reset and post-reset actions) and remove system limitations imposed by the selected mitigation option.

[0035] Method 600 begins in operation 602, when the controller's PDM monitors and detects potential panic situations before they occur. In operation 604, the PCM analyzes the panic situation and implements default mitigation options for the panic situation, including the mitigation options described above. In parallel with operation 604, in operation 606, the PCM proposes several mitigation options to the host device. In operation 608, the controller determines whether a mitigation option decision has been received from the host device. If the controller determines that the host device has not selected a mitigation option, the controller continues to wait for the host device to select one. If the controller determines that the host device has selected a mitigation option, in operation 610, the controller disables the default mitigation option and instead implements the mitigation option selected by the host device. In operation 612, the controller determines whether the panic situation has been resolved. If the panic situation has not been resolved, the controller waits until it is resolved before proceeding to operation 614. In action 614, once the panic situation is resolved, the controller removes the panic instruction and stops performing the mitigation action before returning to action 602.

[0036] Figure 7A is a schematic block diagram showing a memory system 700A for detecting and mitigating future panic situations according to several embodiments. Figure 7B is a flowchart of method 700B for detecting and mitigating panic situations using the memory device of Figure 7A. Since the steps in Figure 7A correspond to the operations of method 700B in Figure 7B, Figure 7A should be read in conjunction with Figure 7B. For example, operation 702B in Figure 7B is associated with step 702A in Figure 7A.

[0037] The storage system 700A comprises a first SSD (e.g., SSD B), a second SSD (e.g., SSD A), an optional switch, a root complex, and host memory (e.g., HMB150 in Figure 1). In some embodiments, one of the mitigation options may be peer-to-peer (P2P) early panic situation handling. In some data centers, the contents of data storage devices (e.g., SSDs) are typically sharded or replicated. Failure (or impending failure) of a data storage device may lead to the host device redirecting input and output (I / O) to another data storage device. Thus, the performance of the storage system may be maintained regardless of failure or degradation of another storage device in the system by adding a status code or other indicator pointing to a secondary location for the requested data.

[0038] However, in some embodiments, if the secondary location is within the same PCIe tree and reachable, the primary data storage device may redirect the command to another data storage device having a PRP / SGL that again points to the same host device region. In some embodiments, the data storage devices may have a submission queue (SQ) between them. The first SSD (e.g., SSD B) may receive the command as is, queue it to the second SSD (e.g., SSD A), and decide to ring the doorbell. The second SSD executes the command and successfully completes it. This method is particularly beneficial for devices that do not use interrupts, such as GPUs, since they cannot currently be moved.

[0039] Method 700B begins with operation 702B, in which the host queues a command in a first data storage device (e.g., SSD B). In operation 704B, the first data storage device detects a potential panic situation before it occurs and detects that another data storage device (e.g., SSD A) can execute the requested command. In operation 706B, the first data storage device queues the modified command in the P2P SQ of the other or second data storage device. In operation 708B, the second data storage device executes the command (e.g., data transfer). In operation 710B, the second data storage device updates the associated P2P completion queue (CQ) and optionally interrupts the first data storage queue. In operation 712B, the first data storage device parses the completion entry. In operation 714B, the first data storage device writes the completion entry to the associated host CQ. In operation 716B, the first data storage device interrupts the host device. Note that although all SQ and CQ are shown in CMD mode in Figure 7A, method 700B can be implemented in host memory.

[0040] Data storage device failures can be mitigated by detecting panic situations and providing the host device with detailed panic data and mitigation options before panic conditions occur. When a future panic situation is detected, several mitigation options may be presented to the host device, such as adjusting read performance, increasing device power, performing evacuation and management actions, and / or redirecting host commands to another data storage device for command completion. By redirecting host commands to another data storage device for command completion, the data storage system becomes more robust, enabling it to handle panic situations with minimal latency impact in the high-end market.

[0041] In one embodiment, the data storage device includes a memory device and a controller coupled to the memory device, the controller being configured to detect a future panic situation of the data storage device, analyze the future panic situation, and propose at least one mitigation option to a host device, wherein the proposal includes a decision timeout; execute a default mitigation option while waiting to receive a selected mitigation option from the host device; and receive a selected mitigation option from the host device, which has been selected from the proposed at least one mitigation option, and execute the selected mitigation option.

[0042] The controller is further configured to execute the default mitigation option based on the decision timeout. Receiving the selected mitigation option overrides the default mitigation option. The controller is not exposed to further future panic situations while waiting to receive the selected mitigation option from the host device. The data storage device and the second data storage device are peer-to-peer (P2P). One of at least one mitigation option redirects host commands received by the data storage device to the second data storage device. Redirecting host commands involves determining whether the second data storage device can execute the host command and queuing the host command in the second data storage device's send queue. The second data storage device's send queue is a peer-to-peer (P2P) send queue. Redirecting host commands further involves parsing a completion entry into the second data storage device's completion queue and writing the completion entry to the data storage device's associated host completion queue. The second data storage device's completion queue is a peer-to-peer (P2P) completion queue. Redirecting host commands further involves writing the completion entry to the relevant host completion queue and then interrupting the host device.

[0043] In another embodiment, the data storage device includes a memory device and a controller coupled to the memory device, the controller being configured to detect a future panic situation of the data storage device based on a panic indicator, analyze the future panic situation, propose at least one mitigation option to the host device, one of which includes redirecting a host command to another location, execute the mitigation option, and determine that the future panic situation has been resolved.

[0044] Another location is a second data storage device. The second data storage device is located within the same PCIe tree as the data storage device. The controller is further configured to remove panic instructions and stop the execution of mitigation options after a future panic situation has been resolved. The controller includes a panic early detection module and a panic control module.

[0045] In another embodiment, the data storage device comprises means for storing data and a controller coupled to the means for storing data, the controller being configured to detect a future panic situation of the data storage device, receive a host command queued in the host device's transmit queue, and, for completion, redirect the host command to a second data storage device, the second data storage device being in the same PCIe tree as the data storage device, and interrupt the host with a completion entry into the host completion queue of the host device.

[0046] The controller is further configured to redirect host commands to a second data storage device using a physical region page (PRP) that points again to the same host region. The controller is further configured to redirect host commands to a second data storage device using a scatter gather list (SGL) that points again to the same host region. The controller is further configured to add an indicator that points to a different location when a host command is received.

[0047] While the above applies to embodiments of the present disclosure, other embodiments and further embodiments of the present disclosure can be devised without departing from the basic scope of the present disclosure, and the scope of the present disclosure is determined by the following claims.

Claims

1. A data storage device, Memory devices and, The memory device is coupled to a controller, and the controller is To detect future panic situations in the aforementioned data storage device, Analyzing the aforementioned future panic situation, To propose at least one mitigation option to the host device, wherein the proposal includes a decision timeout. While waiting for the selected mitigation option to be received from the host device, the default mitigation option is executed. The host device receives the selected relaxation option, which is selected from the proposed at least one relaxation option. A data storage device configured to perform the selected relaxation option described above.

2. The data storage device according to claim 1, wherein the controller is further configured to perform a selected relaxation option based on the decision timeout.

3. The data storage device according to claim 1, wherein receiving the selected relaxation option disables the default relaxation option.

4. The data storage device according to claim 3, wherein the controller is not exposed to further future panic situations while waiting to receive the selected mitigation option from the host device.

5. The data storage device according to claim 1, wherein one of the at least one relaxation option redirects a host command received by the data storage device to a second data storage device.

6. The data storage device according to claim 5, wherein the data storage device and the second data storage device are peer-to-peer (P2P).

7. Redirecting the aforementioned host command Determining whether the second data storage device can execute the host command, The data storage device according to claim 5, further comprising queuing the host command in the transmission queue of the second data storage device.

8. The data storage device according to claim 7, wherein the transmission queue of the second data storage device is a peer-to-peer (P2P) transmission queue.

9. Redirecting the aforementioned host command Analyzing the completion entries to the completion queue of the second data storage device, The data storage device according to claim 7, further comprising writing a completion entry to the associated host completion queue of the data storage device.

10. The data storage device according to claim 9, wherein the completion queue of the second data storage device is a peer-to-peer (P2P) completion queue.

11. The data storage device according to claim 9, further comprising redirecting the host command to interrupt the host device after writing the completion entry to the associated host completion queue.

12. A data storage device, Memory devices and, The memory device is coupled to a controller, and the controller is Based on the panic indicator, the future panic situation of the data storage device is detected. Analyzing the aforementioned future panic situation, The present invention proposes at least one mitigation option to the host device, wherein one of the at least one mitigation option includes redirecting a host command to another location. Implement the mitigation option, A data storage device configured to determine that the aforementioned future panic situation has been resolved.

13. The data storage device according to claim 12, wherein the other location is a second data storage device.

14. The data storage device according to claim 13, wherein the second data storage device is located in the same PCIe tree as the data storage device.

15. The data storage device according to claim 12, wherein the controller is further configured to remove the panic instruction and stop the execution of the mitigation option after the future panic situation has been resolved.

16. The data storage device according to claim 12, wherein the controller comprises a panic early detection module and a panic control module.

17. A data storage device, Means of storing data, The system comprises a controller coupled to means for storing the aforementioned data, and the controller is The data storage device detects future panic situations, The host device receives host commands that have been queued in the send queue. To complete the process, the host command is redirected to a second data storage device, the second data storage device being located in the same PCIe tree as the first data storage device. A data storage device configured to interrupt the host upon completion entries to the host completion queue associated with the host device.

18. The data storage device according to claim 17, wherein the controller is further configured to redirect the host command to the second data storage device using a physical area page (PRP) that again points to the same host area.

19. The data storage device according to claim 17, wherein the controller is further configured to redirect the host command to the second data storage device using a distributed collection list (SGL) that again points to the same host area.

20. The data storage device according to claim 17, wherein the controller is further configured to add an indicator pointing to another location when the host command is received.

Citation Information

Patent Citations

  • Disk array controller

    JP1998078854A

  • Data duplex storage sub-system

    JP1999085410A

  • Transfer method, management device and management program of storage network, and storage network system

    JP2006059119A

  • Power supply unit for storage unit and method for managing storage unit

    JP2007280554A

  • Relocation system and relocation method

    JP2008112276A