HOST-CONTROLLED ATTENTION OF PANIC SITUATIONS

The integration of a panic early warning and control module in data storage devices allows for early detection and mitigation of impending failures, enhancing system reliability by providing proactive host device interventions.

DE102025115501A1Pending Publication Date: 2026-03-05SANDISK TECHNOLOGIES LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE102025115501
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Current data storage devices lack the capability to prepare for and mitigate panic situations in advance, leading to potential failures and a lack of effective mitigation strategies when such situations occur.

Method used

Implementing a panic early warning module and control module in data storage devices to detect impending panic situations, providing detailed panic data and mitigation options to the host device, allowing for proactive adjustments such as redirecting commands, adjusting performance, or swapping operations to other storage devices.

Benefits of technology

Reduces data storage device failures by enabling early detection and mitigation of panic situations, maintaining system performance and reliability through proactive host device interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

By detecting panic situations and providing detailed panic data and mitigation options to the host device before the panic situation occurs, data storage device failure can be reduced. Upon detection of an impending panic situation, the host device can be offered several mitigation options, such as adjusting read performance; increasing device performance; performing swapping and management operations; and / or redirecting host commands to another data storage device for command execution. If the other data storage device is in the same PCIe tree and reachable, the command can be redirected using PRPs / SGLs that point back to the same host device. In some embodiments, the data storage devices can have a transmission queue between them.In some embodiments, the data storage device includes a panic early warning module and a panic control module.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF REVELATION Area of ​​Revelation

[0001] Embodiments of the present disclosure generally relate to a data storage device for the early detection and mitigation of panic situations. Description of the state of the art

[0002] A device panic situation, or panic situation, refers to circumstances in which a data storage device (e.g., a solid-state drive (SSD)) can notify a host device of mitigation measures to be taken when a panic condition (e.g., a failure) occurs. The panic condition can be signaled by an asynchronous event or a controller failure status register. Once a panic condition, such as a failure, occurs, the host device can perform a reset action. After a host device has performed mitigation measures in accordance with the detected panic condition, the storage device is expected to provide diagnostic information. Under certain circumstances, the data storage device may also suggest (potential) post-reset actions that should be taken.

[0003] When a data storage device detects a panic situation, the panic situation is reported to the host device via an interface, and the data storage device's capabilities during the panic situation are made available to the host device. However, there are currently no requirements or procedures for a data storage device to prepare for a future panic situation in advance of its occurrence.

[0004] Therefore, there is a need for an improved data storage device for the early detection and mitigation of panic situations. SUMMARY OF THE REVELATION

[0005] By detecting panic situations and providing detailed panic data and mitigation options to the host device before the panic situation occurs, data storage device failure can be reduced. Upon detection of an impending panic situation, the host device can be offered several mitigation options, such as adjusting read performance; increasing device performance; performing swapping and management operations; and / or redirecting host commands to another data storage device for command execution. If the other data storage device is in the same PCIe tree and reachable, the command can be redirected using PRPs / SGLs that point back to the same host device. In some embodiments, the data storage devices can have a transmission queue between them.In some embodiments, the data storage device includes a panic early warning module and a panic control module.

[0006] In one embodiment, a data storage device includes a storage device and a controller coupled to the storage device, the controller being configured to: detect a future panic situation of the data storage device; analyze the future panic situation; propose at least one mitigation option to a host device, the proposal including a decision timeout; execute a default mitigation option while waiting to receive the selected mitigation option from the host device; receive a selected mitigation option from the host device, chosen from the at least one proposed mitigation option; and execute the selected mitigation option.

[0007] In another embodiment, a data storage device includes a storage device and a controller coupled to the storage device, the controller being configured to: detect a future panic situation of the data storage device based on a panic indicator; analyze the future panic situation; propose at least one mitigation option to a host device, wherein one of the at least mitigation options includes redirecting a host command to another location; execute a mitigation option; and determine that the future panic situation is resolved.

[0008] In yet another embodiment, the data storage device includes means for storing data; and a controller coupled to the means for storing data, wherein the controller is configured to: detect a future panic situation of the data storage device; receive a host command that is queued in a transmission queue of a host device; redirect the host command for completion to a second data storage device, wherein the second data storage device is in the same PCIe tree as the data storage device; and interrupt the host with a termination entry in a corresponding host termination queue of the host device. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to understand in detail the nature of the aforementioned features of the present disclosure, a more precise description of the disclosure, which has been briefly summarized above, can be given with reference to embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings only illustrate typical embodiments of this disclosure and are therefore not to be considered as limiting in its scope of protection, since the disclosure also permits other, equally effective embodiments. Fig. Figure 1 is a schematic block diagram illustrating a storage system in which a data storage device can function as a storage device for a host device according to certain embodiments. Fig. Figure 2 is a table illustrating various panic reset and recovery actions of a data storage device according to some embodiments. Fig. Figure 3 is a schematic block diagram illustrating a storage system with panic early detection and control according to some embodiments. Fig. Figure 4 is a table illustrating various potential fault injection types for resolving panic situations in a data storage device according to some embodiments. Fig. Figure 5 is a flowchart illustrating a method for panic situation detection and mitigation in a data storage device according to some embodiments. Fig. Figure 6 is a flowchart illustrating a method for panic situation detection and mitigation in a data storage device according to some embodiments. Fig. Figure 7A is a schematic block diagram illustrating a memory system for detecting and mitigating future panic situations according to some embodiments. Fig. 7B is a flowchart that describes a procedure for panic situation detection and mitigation in the storage device of Fig. 7A illustrates.

[0010] To facilitate understanding, identical reference numerals have been used wherever possible to denote identical elements present in all figures. It is intended that elements disclosed in one embodiment may also be used effectively in other embodiments without specific mention. DETAILED DESCRIPTION

[0011] The following refers to embodiments of the disclosure. It is understood, however, that the disclosure is not limited to the specific embodiments described. Instead, any combination of the following features and elements, regardless of whether they relate to different embodiments or not, is intended for the implementation and practical application of the disclosure. Furthermore, although embodiments of the disclosure may offer advantages over other possible solutions and / or over the prior art, the fact that a particular embodiment achieves a specific advantage or not does not constitute a limitation of the disclosure.Therefore, the following aspects, features, embodiments, and advantages serve only for illustration and are not considered elements or limitations of the appended claims unless expressly stated in one or more claims. Likewise, a reference to "the disclosure" is not to be construed as a generalization of any inventive subject matter disclosed herein and is not to be considered an element or limitation of the appended claims unless expressly stated in one or more claims.

[0012] By detecting panic situations and providing detailed panic data and mitigation options to the host device before the panic situation occurs, data storage device failure can be reduced. Upon detection of an impending panic situation, the host device can be offered several mitigation options, such as adjusting read performance; increasing device performance; performing swapping and management operations; and / or redirecting host commands to another data storage device for command execution. If the other data storage device is in the same PCIe tree and reachable, the command can be redirected using PRPs / SGLs that point back to the same host device. In some embodiments, the data storage devices can have a transmission queue between them.In some embodiments, the data storage device includes a panic early warning module and a panic control module.

[0013] Fig. Figure 1 is a schematic block diagram illustrating a storage system 100 comprising a data storage device 106, which, according to certain embodiments, can function as a storage device for a host device 104. For example, the host device 104 can use non-volatile memory (NVM) 110 enclosed in the data storage device 106 for storing and retrieving data. The host device 104 includes dynamic random-access memory (DRAM) 138. In some examples, the storage system 100 can include a variety of storage devices, such as the data storage device 106, which can function as a storage array.For example, the storage system 100 can include a variety of data storage devices 106 configured as a redundant array of inexpensive / independent disks (RAID) and acting collectively as a mass storage device for the host device 104.

[0014] The host device 104 can store data on and / or retrieve data from one or more storage devices, such as the data storage device 106. As described in Fig. As illustrated in Figure 1, the host device 104 can communicate with the data storage device 106 via an interface 114. The host device 104 can include a wide range of devices, including computer servers, network-attached storage (NAS) units, desktop computers, notebooks (i.e., laptops), tablet computers, set-top boxes, mobile phones such as smartphones or smart tablets, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, or other devices that can send data to or receive data from a data storage device.

[0015] The host DRAM 138 can optionally include a host memory buffer (HMB) 150. The HMB 150 is a section of the host DRAM 138 allocated to the data storage device 106 for the exclusive use of a controller 108 of the data storage device 106. For example, the controller 108 can store mapping data, buffered instructions, logic-physical (L2P) tables, metadata, and the like in the HMB 150. In other words, the HMB 150 can be used by the controller 108 to store data that would normally be stored in volatile memory 112, a buffer 116, internal memory of the controller 108 such as static random-access memory (SRAM), and the like. In examples where the data storage device 106 does not include DRAM (i.e., optional DRAM 118), the controller 108 can use the HMB 150 as the DRAM of the data storage device 106.

[0016] The data storage device 106 includes a controller 108, an NVM 110, a power supply 111, volatile memory 112, an interface 114, a write buffer 116, and an optional DRAM 118. In some examples, the data storage device 106 may include additional components, which for clarity are shown in Fig. Figure 1 is not shown. For example, the data storage device 106 may include a printed circuit board (PCB) to which components of the data storage device 106 are mechanically attached and which includes electrically conductive traces that electrically connect components of the data storage device 106 or the like. In some examples, the physical dimensions and connection configurations of the data storage device 106 may conform to one or more standard form factors. Some examples of standard form factors include, but are not limited to, 3.5-inch data storage devices (e.g., an HDD or SSD), 2.5-inch data storage devices, 1.8-inch data storage devices, Peripheral Component Interconnect (PCI), PCI-Extended (PCI-X), and PCI Express (PCIe) (e.g., PCIe x1, x4, x8, x16, PCIe Mini Card, MiniPCI, etc.).In some examples, the data storage device 106 can be directly coupled to a mainboard of the host device 104 (e.g., directly soldered or plugged into a connector).

[0017] Interface 114 can include a data bus for data exchange with the host device 104 and / or a control bus for exchanging commands with the host device 104. Interface 114 can operate according to any suitable protocol. For example, interface 114 can operate according to the Non-Volatile Memory Express (NVMe) protocol or the like. Interface 114 (e.g., the data bus, the control bus, or both) is electrically connected to the controller 108 and provides an electrical connection between the host device 104 and the controller 108, allowing data to be exchanged between them. In some examples, the electrical connection of interface 114 can also allow the data storage device 106 to receive power from the host device 104. For example, in Fig. As illustrated in Figure 1, the power supply 111 can receive power from the host device 104 via the interface 114.

[0018] The NVM 110 can include a variety of storage devices or storage units. The NVM 110 can be configured to store and / or retrieve data. For example, a storage unit of the NVM 110 can receive data and a message from the controller 108 instructing the storage unit to store the data. Likewise, the storage unit can receive a message from the controller 108 instructing the storage unit to retrieve data. In some examples, each of the storage units can be referred to as a chip. In some examples, the NVM 110 can include a variety of chips (i.e., a variety of storage units). In some examples, each storage unit can be configured to store relatively large amounts of data (e.g., 128 MB, 256 MB, 512 MB, 1 GB, 2 GB, 4 GB, 8 GB, 16 GB, 32 GB, 64 GB, 128 GB, 256 GB, 512 GB, 1 TB, etc.).

[0019] In some examples, each storage unit can include any type of non-volatile storage device, such as flash memory devices, phase-change memory (PCM), resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), holographic storage devices, and any other type of non-volatile storage device.

[0020] The NVM 110 can comprise a variety of flash memory devices or storage units. NVM flash memory devices can include NAND- or NOR-based flash memory devices and can store data based on a charge contained in a floating gate of a transistor for each flash memory cell. In NVM flash memory devices, the flash memory device can be subdivided into a multitude of chips, with each chip containing a multitude of physical or logical blocks, which can be further subdivided into a multitude of pages. Each block within a given memory device can contain a multitude of NVM cells. Rows of NVM cells can be electrically connected using a word line to define a page within a multitude of pages.The individual cells in each of the multitude of pages can be electrically connected to the respective bit lines. Furthermore, NVM flash storage devices can be 2D or 3D devices and can be of the single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), or quad-level cell (QLC) type. The controller 108 can write and read data to and from NVM flash storage devices at the page level and erase data from NVM flash storage devices at the block level.

[0021] The power supply 111 can power one or more components of the data storage device 106. In standard mode, the power supply 111 can power one or more components by drawing power from an external device, such as the host device 104. For example, the power supply 111 can power the one or more components by using power received from the host device 104 via interface 114. In some examples, the power supply 111 can include one or more power storage components configured to power the one or more components when they are in shutdown mode, such as when no power is being received from the external device. In this way, the power supply 111 can function as an integrated backup power source.Examples of one or more energy storage components include capacitors, supercapacitors, batteries, and the like. In some cases, the amount of electricity that can be stored by one or more energy storage components may depend on the cost and / or size (e.g., area / volume) of the one or more energy storage components. In other words, as the amount of electricity stored by one or more energy storage components increases, so do the cost and / or size of the one or more energy storage components.

[0022] The volatile memory 112 can be used by the controller 108 to store information. The volatile memory 112 can include one or more volatile storage devices. In some examples, the controller 108 can use the volatile memory 112 as a cache. For example, the controller 108 can store cached information in the volatile memory 112 until the cached information is written to the NVM 110. As in Fig. As illustrated in Figure 1, the volatile memory 112 can consume power received from the power supply 111. Examples of volatile memory 112 include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static RAM (SRAM), and synchronous dynamic RAM (SDRAM (e.g., DDR1, DDR2, DDR3, DDR3L, LPDDR3, DDR4, LPDDR4, and the like)). Likewise, the optional DRAM 118 can be used to store mapping data, buffered instructions, logic-physical (L2P) tables, metadata, cached data, and the like. In some examples, the data storage device 106 does not include the optional DRAM 118, so the data storage device 106 has no DRAM. In other examples, the data storage device 106 includes the optional DRAM 118.

[0023] The controller 108 can manage one or more operations of the data storage device 106. For example, the controller 108 can manage reading data from and / or writing data to the NVM 110. In some embodiments, when the data storage device 106 receives a write command from the host device 104, the controller 108 can initiate a data storage command to store data in the NVM 110 and monitor the progress of the data storage command. The controller 108 can determine at least one operating property of the storage system 100 and store at least one operating property in the NVM 110. In some embodiments, when the data storage device 106 receives a write command from the host device 104, the controller 108 temporarily stores the data associated with the write command in internal memory or write buffer 116 before sending the data to the NVM 110.The controller 108 can include switching logic or processors configured to run programs to operate the data storage device 106.

[0024] The controller 108 can include an optional second volatile memory 120. The optional second volatile memory 120 can be similar to the volatile memory 112. For example, the optional second volatile memory 120 can be SRAM. The controller 108 can allocate a portion of the optional second volatile memory to the host device 104 as a control memory buffer (CMB) 122. The CMB 122 can be accessed directly by the host device 104. For example, instead of managing one or more transmission queues in the host device 104, the host device 104 can use the CMB 122 to store the one or more transmission queues that are normally managed in the host device 104.In other words, the host device 104 can generate commands and store the generated commands with or without the associated data in the CMB 122, with the controller 108 accessing the CMB 122 to retrieve the stored generated commands and / or the associated data.

[0025] Fig. Figure 2 is Table 200, which illustrates various panic reset and recovery actions of a data storage device according to some embodiments. Table 200 is taken from the publicly available Open Compute Project (OCP) datacenter specification titled Datacenter NVMe® SSD Specification (Version 2.0). A data storage device may use a bit field to indicate potential reset actions that must be performed during or before a panic situation to prevent it. The data storage device may also use a bit field to specify an appropriate device recovery action for handling a panic situation (e.g., device panic state or panic mode).As discussed below, the probability of device failure can be reduced by providing the host device with additional mitigation options and detailed panic data before the device panic state occurs (where possible), and by providing the host device with a set of recovery / mitigation options to handle the device panic state. The set of mitigation options may include modifying a subset of various device capabilities, such as reducing power, increasing power drawn, removing the option to read from a subset of the chips, or even redirecting host commands to another available storage device.In some embodiments, such a storage system analyzes the current situation based on inputs regarding the state of the storage device and external conditions, and provides an indication of the potential panic state to the host using a log page or other suitable means of signaling.

[0026] In some embodiments, if the data storage device provides the host device with a set of recovery / mitigation options to handle the device panic state, it may have reduced device capabilities for the duration of the device panic state. It should be noted that while the failure recovery protocol is designed for data center storage devices (e.g., enterprise storage devices), the protocol could also be adapted for client SSDs. The disclosed embodiments are applicable to both types of SSDs, data center storage devices and client SSDs, although certain characteristics may vary. For example, capacitor or DRAM failures may not be applicable to client SSDs, but HMB failures may occur in client devices but not in data center devices.

[0027] Fig. Figure 3 is a schematic block diagram illustrating a storage system 300 with panic early detection and control according to some embodiments. The storage system 300 comprises a host device 302, a data storage controller 304, and NVM chips 310. The host device 302 can be the host 104 of Fig. 1. The data storage controller 304 can control the controller 108 from Fig. 1. The NVM 310 chips can be used in the NVM 110 from Fig. 1. The data storage controller 304 includes a panic early warning module (PDM) 306 and a panic control module (PCM) 308.

[0028] The PDM 306 is configured to detect a panic situation. The goal is to detect a future panic situation as early as possible. Consequently, the system anticipates a certain false alarm rate if it detects a future panic situation but it is averted. The PCM 308 is configured to provide the host device 302 with several mitigation options to modify the storage system behavior once the PDM 306 detects a future device panic state, depending on the panic ID (e.g., the fault type injection from the system). Fig. 4).

[0029] Fig. Figure 4 is Table 400, which illustrates various potential fault injection types for resolving panic situations of a data storage device according to some embodiments. Table 400 is taken from the OCP Datacenter Specification. The possible causes of device panic states range from firmware failure to NVM failure. The panic IDs and their associated causes vary; examples of different possible fault injection types for remediation are shown in Table 400. Thus, if the data storage device (e.g., Data Storage Device 304 of Fig. 3) If a panic situation is detected and a panic ID is determined that relates to the cause, the data storage device can assign the panic situation, the panic situation failure conditions, or the panic ID to the appropriate fault injection type in Table 400 for remediation.

[0030] Device panic situations, such as hardware malfunctions, can be associated with a wide range of characteristics. While the detection of these characteristics by the data storage device may indicate a potential panic situation, not all characteristics—or any single characteristic alone—necessarily lead to a panic situation or claim. For example, in the case of a malfunction in an error correction code engine (ECC engine), the malfunction may be detected early via a PDM (e.g., PDM 306). Fig. 3) can be detected by observing a drop in performance or an increase in power consumption. Under this condition, the host device can be offered a choice. For example, without triggering a shutdown, the host device can decide whether to maintain reduced read performance at the same power consumption or to consume more power while maintaining the same performance.

[0031] In another example, if a NAND chip fails, one of the chips may become unreadable. This can be detected early through monitoring if the number of write cycles required is unusually high. Under these circumstances, the system would likely swap all data to other chips. However, this represents a trade-off for the host device. The host device may experience reduced write performance but can still read from the storage device, albeit with some reduced reliability, until the faulty chip eventually becomes unreadable. Alternatively, the host device can provide the data storage device with time windows to perform management operations, such as swapping and maintenance operations, and attempts to revive the chip for write commands.After this period, the host device can eliminate the reduced reliability problems, so that in situations where the chip has not been successfully revived, only the reduced write performance remains.

[0032] In yet another example, in some cases a portion of the DRAM may be considered damaged. Consequently, the DRAM can employ a special ECC (Emergency Compatibility Check), and as soon as a problem is detected, the storage device can display a panic situation early warning flag (i.e., before actual damage is detected). Once the extent of the damaged DRAM is determined, the controller (e.g., controller 108 of Fig. 1) Suggest reducing the power or exported capacity so that the controller has to "control" less data with the remaining DRAM. The controller may also suggest disabling features that rely on DRAM or utilizing most of the host device's DRAM (e.g., the host device's HMB, such as the HMB 150). Fig. 1) if available. Alternatively, the controller may decide to increase the ECC bits while operating with DRAM to increase integrity. In some embodiments, the data storage device's decisions should be relayed to the controller within a specified timeframe via the same interface; otherwise, the data storage device will choose a default option to mitigate the detected early panic state.

[0033] Fig. Figure 5 is a flowchart illustrating a method 500 for panic situation detection and mitigation in a data storage device according to some embodiments. The method 500 begins with operation 502, in which a PDM (e.g., PDM 306 of Fig. 3) a controller (e.g., controller 108 of Fig. 1) It monitors and detects a potential panic situation before it occurs. In Operation 504, the PCM (e.g., PCM 308 from) analyzes Fig. 3) the panic situation and proposes several mitigation options for the panic situation, such as the mitigation options discussed above. In Operation 506, the early panic situation (e.g., the cause of the panic situation and other information about the panic situation) and the mitigation options are sent to the host device with a timeout for the host device to make a decision. In Operation 508, the data storage device determines whether the host device has decided on mitigation options within the specified timeout. If the host device has decided on a mitigation option within the specified timeout by informing the controller of the chosen mitigation, in Operation 510 the controller executes the changes specified by the chosen mitigation option.If the attenuation options sent to the host device time out and the host device has not selected an attenuation option, the controller performs the changes specified by a default attenuation option during operation 512.

[0034] Fig. Figure 6 is a flowchart illustrating a method 600 for panic situation detection and mitigation in a data storage device according to some embodiments. In some embodiments, the controller (e.g., the controller 108 of Fig. 1) if a PDM (e.g. PDM 306 from Fig. 3) recognizes a potential panic situation before the panic situation occurs, and a PCM (e.g., PCM 308 from Fig. 3) Analyzes the available and appropriate mitigation options for the detected panic situation and immediately switches to a default mitigation option from among the various appropriate options. In parallel, the corresponding mitigation options are also sent to the host device and announced, where the host device can override the implemented default mitigation option by selecting a mitigation option. This prevents the data storage device from experiencing further failures while the host selects a mitigation option, which in turn improves the device's condition. In some embodiments, the panic situation may be temporary, and the controller manages to resolve it. Under these circumstances, the controller can use the interface (e.g., interface 114 of Fig. 1) Use to remove the panic indicator (which may sometimes require reset and post-reset operations) and to remove the system limitation imposed by the selected mitigation option.

[0035] Procedure 600 begins with Operation 602, in which a controller's PDM monitors for and detects a potential panic situation before it occurs. In Operation 604, the PCM analyzes the panic situation and implements a standard panic mitigation option, such as the mitigation options discussed above. Concurrently with Operation 604, in Operation 606, the PCM proposes several mitigation options to the host device. In Operation 608, the controller determines whether a decision regarding a mitigation option has been received from the host device. If the controller determines that the host device has not selected a mitigation option, the controller continues to wait for the host device to select one.If the controller determines that the host device has selected an attenuation option, in operation 610 the controller overrides the default attenuation option and instead executes the attenuation option selected by the host device. In operation 612, the controller determines whether the panic situation has been resolved. If the panic situation has not been resolved, the controller waits until it is before proceeding to operation 614. Once the panic situation is resolved, in operation 614 the controller removes the panic alert and terminates the attenuation operations before returning to operation 602.

[0036] Fig. Figure 7A is a schematic block diagram illustrating a 700A storage system for detecting and mitigating future panic situations according to some embodiments. Fig. 7B is a flowchart that describes a procedure 700B for panic situation detection and mitigation in the storage device of Fig. 7A illustrates. Fig. 7A is in conjunction with Fig. 7B to read, as the steps of Fig. 7A the operations of procedure 700B of Fig. 7B corresponds. For example, operation 702B is the Fig. 7B with step 702A of the Fig. 7A linked etc.

[0037] The 700A storage system includes a first SSD (e.g., SSD B), a second SSD (e.g., SSD A), an optional switch, a root complex, and a host memory (e.g., HMB 150). Fig. 1) In some embodiments, one of the mitigation options may be handling an early panic situation through peer-to-peer (P2P) communication. The contents of the data storage device (e.g., SSDs) are commonly shared or duplicated in some data centers. The failure (or imminent failure) of data storage devices may cause the host device to redirect input / output (I / O) traffic to another data storage device. This allows the performance of the storage system to be maintained despite the failure or reduced performance of another storage device in the system by adding a status code or other indicator that points to a secondary location for the requested data.

[0038] In some embodiments, however, a primary data storage device can redirect the command to another data storage device using PRPs / SGLs that point back to the same region of the host device, provided the secondary location is in the same PCIe tree and reachable. In some embodiments, the data storage devices can have a submission queue (SQ) between them. A first SSD (e.g., SSD B) can decide to accept a command unchanged, place it in the queue of a second SSD (e.g., SSD A), and ring the doorbell. The second SSD executes the command and completes it normally. This method is particularly advantageous for devices that do not use interrupts, such as GPUs, since these cannot currently be moved.

[0039] Procedure 700B begins with operation 702B, in which the host queues a command on a first storage device (e.g., SSD B). In operation 704B, the first storage device detects a potential panic situation before it occurs and that another storage device (e.g., SSD A) can execute the requested command. In operation 706B, the first storage device queues a revised command in the peer-to-peer (P2P) queue of the other, or second, storage device. In operation 708B, the second storage device executes the command (e.g., data transfer). In operation 710B, the second storage device updates the appropriate P2P completion queue (CQ) and optionally suspends the first storage device's queue. In operation 712B, the first storage device parses the completion entry.In operation 714B, the first data storage device writes a completion entry to the relevant host CQ. In operation 716B, the first data storage device interrupts the host device. Note that in . Fig. 7A shows all SQs and CQs in CMD mode, but procedure 700B could also be implemented in host memory.

[0040] By detecting panic situations and providing detailed panic data and mitigation options to the host device before the panic situation occurs, data storage device failure can be reduced. Upon detection of an impending panic situation, the host device can be offered several mitigation options, such as adjusting read performance; increasing device performance; performing offloading and management operations; and / or redirecting host commands to another data storage device for command execution. Redirecting host commands to another data storage device for command completion makes the data storage system more robust, as it is able to handle panic situations with minimal latency impact in a high-end market.

[0041] In one embodiment, a data storage device includes a storage device and a controller coupled to the storage device, the controller being configured to: detect a future panic situation of the data storage device; analyze the future panic situation; propose at least one mitigation option to a host device, the proposal including a decision timeout; execute a default mitigation option while waiting to receive a selected mitigation option from the host device; receive a selected mitigation option from the host device, chosen from the at least one proposed mitigation option; and execute the selected mitigation option.

[0042] The controller is further configured to execute a default attenuation option based on the decision timeout. Receiving the selected attenuation option overrides the default attenuation option. While waiting to receive the selected attenuation option from the host device, the controller is not exposed to any further future panic situations. The data storage device and the second data storage device are peer-to-peer (P2P). One of the at least two attenuation options redirects a host command received from the data storage device to a second data storage device. Redirecting the host command includes: determining whether the second data storage device can execute the host command; and queuing the host command in a delivery queue on the second data storage device.The second data storage device's delivery queue is a peer-to-peer delivery queue (P2P delivery queue). Redirecting the host command further includes: parsing a completion entry into a completion queue of the second data storage device; and writing a completion entry to a relevant host completion queue of the data storage device. The second data storage device's completion queue is a peer-to-peer completion queue (P2P completion queue). Redirecting the host command further includes suspending the host device after writing the completion entry to the relevant host completion queue.

[0043] In another embodiment, a data storage device includes a storage device and a controller coupled to the storage device, the controller being configured to: detect a future panic situation of the data storage device based on a panic indicator; analyze the future panic situation; propose at least one mitigation option to a host device, wherein one of the at least mitigation options includes redirecting a host command to another location; execute a mitigation option; and determine that the future panic situation is resolved.

[0044] The other location is a second data storage device. The second data storage device is located in the same PCIe tree as the primary data storage device. The controller is further configured to remove the panic alert and stop executing the mitigation option after the future panic situation has resolved. The controller includes a panic early warning module and a panic control module.

[0045] In yet another embodiment, the data storage device includes means for storing data; and a controller coupled to the means for storing data, wherein the controller is configured to: detect a future panic situation of the data storage device; receive a host command that is queued in a transmission queue of a host device; redirect the host command for completion to a second data storage device, wherein the second data storage device is in the same PCIe tree as the data storage device; and interrupt the host with a termination entry in a corresponding host termination queue of the host device.

[0046] The controller is further configured to redirect the host command to the second data storage device using Physical Region Pages (PRPs) that point back to the same host region. The controller is further configured to redirect the host command to the second data storage device using Scatter-Gather Lists (SGLs) that point back to the same host region. The controller is further configured to add an indicator pointing to a different location when the host command is received.

[0047] While the foregoing relates to embodiments of the present disclosure, other and further embodiments of the disclosure may be conceived without deviating from its basic scope, the scope of which is determined by the following claims.

Claims

[1] Data storage device comprising: a storage device; and a controller coupled to the storage device, wherein the controller is configured to: Detecting a future panic situation of the data storage device; Analyzing the future panic situation; Proposing at least one mitigation option for a host device, the proposal including a decision timeout; Execute a standard attenuation option while waiting to receive a selected attenuation option from the host device; Receiving the selected attenuation option from the host device, which was chosen from the at least one proposed attenuation option; and Execute the selected attenuation option. [2] Data storage device according to claim 1, wherein the controller is further configured to execute a selected attenuation option based on the decision timeout. [3] Data storage device according to claim 1, wherein receiving the selected attenuation option overrides the default attenuation option. [4] Data storage device according to claim 3, wherein the controller is not subject to any further future panic situations while waiting to receive the selected attenuation option from the host device. [5] Data storage device according to claim 1, wherein an attenuation option of the at least one attenuation option redirects a host command received by the data storage device to a second data storage device. [6] Data storage device according to claim 5, wherein the data storage device and the second data storage device are peer-to-peer (P2P). [7] Data storage device according to claim 5, comprising redirecting the host command: Determine whether the second data storage device can execute the host command; and Inserting the host command into a delivery queue on the second data storage device. [8] Data storage device according to claim 7, wherein the transmission queue of the second data storage device is a peer-to-peer transmission queue (P2P transmission queue). [9] Data storage device according to claim 7, wherein the redirection of the host command further comprises: Parsing a completion entry into a completion queue of the second data storage device; and Writing a completion entry to a relevant host completion queue of the data storage device. [10] Data storage device according to claim 9, wherein the completion queue of the second data storage device is a peer-to-peer completion queue (P2P completion queue). [11] Data storage device according to claim 9, wherein redirecting the host command further comprises interrupting the host device after writing the completion entry to the relevant host completion queue. [12] Data storage device comprising: a storage device; and a controller coupled to the storage device, wherein the controller is configured to: Recognizing a future panic situation Data storage device based on a panic indicator; Analyzing the future panic situation; Proposing at least one mitigation option for a host device, wherein one of the at least one mitigation options involves redirecting a host command to another location; Executing a mitigation option; and Determine that the future panic situation has been resolved. [13] Data storage device according to claim 12, wherein the other location is a second data storage device. [14] Data storage device according to claim 13, wherein the second data storage device is located in the same PCIe tree as the data storage device. [15] Data storage device according to claim 12, wherein the controller is further configured to remove the panic alert and to stop executing the mitigation option after the future panic situation has been resolved. [16] Data storage device according to claim 12, wherein the control unit comprises a panic early detection module and a panic control module. [17] Data storage device comprising: Means for storing data; and a control system that is coupled with the means for storing data, where the controller is configured to: Detecting a future panic situation of the data storage device; Receiving a host command that is queued in a host device's transmission queue; Redirecting the host command for completion to a second data storage device, wherein the second data storage device is located in the same PCIe tree as the data storage device; and Interrupting the host with a termination entry in the appropriate host termination queue of the host device. [18] Data storage device according to claim 17, wherein the controller is further configured to redirect the host command to the second data storage device with physical region pages (PRPs) that refer back to the same host region. [19] Data storage device according to claim 17, wherein the controller is further configured to redirect the host command to the second data storage device containing scatter-gather lists (SGLs) that refer back to the same host region. [20] Data storage device according to claim 17, wherein the controller is further configured to add an indicator that points to a different location when the host command is received.