Resume mode for memory device
By introducing recovery boot mode and backup firmware in nonvolatile memory devices, communication interruption problems caused by failures are solved, on-site recovery of failures and efficient utilization of resources are achieved, and overall replacement and analysis time is reduced.
Patent Information
- Application Number
- CN202510082582.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-26
- Filing Date
- 2025-01-20
- Publication Date
- 2025-07-25
AI Technical Summary
Nonvolatile memory devices cannot communicate with the host system in the event of a failure, resulting in waste of resources and power, and the prior art requires the overall replacement of PCBs for analysis, increasing costs and time.
The memory device initializes the recovery boot mode by receiving the fault indication, reboots using backup firmware, and transmits status information to resolve the fault and avoid waste of resources and power.
It realizes on-site recovery of memory devices in the event of a failure, reduces the physical damage and analysis time of the PCB, and improves system performance and resource utilization efficiency.
Smart Images

Figure CN120375893A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This patent application claims priority to U.S. Provisional Patent Application No. 63 / 624,571, filed on January 24, 2024, and titled "Recovery Mode for Memory Device". The disclosure of the prior application is considered part of this patent application and is incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to memory devices, memory device operations, and, for example, recovery modes of memory devices. Background Art
[0004] Non-volatile memory devices, such as NAND memory devices, may use circuitry to enable electrical programming, erasing, and storing of data even when power is not supplied. Non-volatile memory devices may be used in various types of electronic devices, such as computers, mobile phones, or automotive computing systems, among other examples.
[0005] A non-volatile memory device may include a memory cell array, page buffers, and column decoders. Additionally, a non-volatile memory device may include a control logic unit (e.g., a controller), a row decoder, or an address buffer, among other examples. The memory cell array may include memory cell strings connected to bit lines that extend in a column direction.
[0006] Memory cells of a non-volatile memory device (which may be referred to as "cells" or "data cells") may include a current path formed between a source and a drain on a semiconductor substrate. The memory cells may further include a floating gate and a control gate formed between insulating layers on the semiconductor substrate. A programming operation of the memory cells (sometimes referred to as a write operation) is typically achieved by grounding the source and drain regions of the memory cells and the semiconductor substrate of the body region and applying a high positive voltage (which may be referred to as a "programming voltage", "programming power supply voltage", or "VPP") to the control gate to create Fowler-Nordheim (F-N) tunneling (referred to as "F-N tunneling") between the floating gate and the semiconductor substrate. When F-N tunneling occurs, electrons in the body region accumulate on the floating gate through the electric field of the VPP applied to the control gate to increase the threshold voltage of the memory cells.
[0007] An erase operation of memory cells is concurrently performed in units of sectors (referred to as "blocks") of the body region by applying a high negative voltage (which may be referred to as an "erase voltage" or "Vera") to the control gate and applying a configured voltage to the body region to generate F-N tunneling. In this case, electrons accumulated on the floating gate are discharged into the source region, such that the memory cells have an erase threshold voltage distribution.
[0008] Each memory cell string may have a plurality of floating gate type memory cells connected in series with each other. Access lines (sometimes referred to as "word lines") extend in the row direction, and the control gate of each memory cell is connected to a corresponding access line. The nonvolatile memory device may include a plurality of page buffers connected between the bit lines and the column decoder. The column decoder is connected between the page buffer and the data line. SUMMARY OF THE INVENTION
[0009] One aspect of the present disclosure discloses a memory device, including: one or more components configured to: receive an indication that the memory device is associated with a failure from a host system; receive a request to initialize a recovery of the memory device in response to the failure from the host system; initialize a reboot of the memory device based on the request; and transmit status information obtained when rebooting the memory device, wherein the status information includes information associated with a current operating state of the memory device.
[0010] Another aspect of the present disclosure discloses a method, including: receiving, by a memory device, a request to initialize a recovery of the memory device from a host system, wherein the request is based on a failure associated with the memory device; initializing a reboot of the memory device based on the request; and transmitting status information obtained when rebooting the memory device, wherein the status information includes information associated with a current operating state of the memory device.
[0011] Another aspect of the present disclosure discloses a system, including: a host system configured to: detect a failure associated with a memory device, and transmit a request to the memory device to initialize a recovery of the memory device in response to the failure; and the memory device configured to: receive the request from the host system, initialize a reboot of the memory device based on the request, and transmit status information obtained when rebooting the memory device, wherein the status information includes information associated with a current operating state of the memory device. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a diagram illustrating an example system capable of implementing a recovery mode of a memory device.
[0013] Figure 2 It is a diagram illustrating an example of implementing a recovery mode of a memory device.
[0014] Figure 3 It is a diagram illustrating an example of implementing a recovery mode of a memory device.
[0015] Figure 4 It is a flowchart of an example method associated with implementing a recovery mode of a memory device. Detailed Description
[0016] A memory device (e.g., a non-volatile memory (NVM) solid state drive (SSD) associated with a vehicle) can receive commands from a host system to read data from and / or write data to the memory device. In some cases, the memory device can experience a fault. The fault can cause the memory device to be unable to receive and / or process commands. The fault can cause the memory device to no longer communicate with the host system.
[0017] As an example, the SSD may fail to boot in the host system, which can occur when the SSD has a successful peripheral component interconnect express (PCIe) link but the NVM subsystem initialization fails. The failure to boot in the host system can lead to unpredictable host behavior because the host system may attempt to interrogate the SSD but fail. As another example, an embedded or automotive SSD in a ball grid array (BGA) form factor may experience a fault where the SSD firmware can no longer communicate with the host system. A critical automotive SSD fault can be attributed to the loss of communication with the host system (e.g., the memory device may no longer function or may otherwise not respond).
[0018] When the memory device is associated with a fault (e.g., the memory device does not respond to the host system), the host system may waste resources and / or power when attempting to communicate with the non-responsive memory device. The memory device may also waste power because the memory device can be powered on but may be unable to perform any functions.
[0019] When a memory device is associated with a fault, the host system associated with the memory device may be recalled to a service center because no action can be taken on-site. When the host system and the memory device are associated with a vehicle, the entire vehicle may be recalled to a service center, which may frustrate the end customer (e.g., the vehicle owner). At the service center, the printed circuit board (PCB) with the SSD is typically replaced as a whole and then sent for further analysis because the SSD is usually a BGA and soldered onto the PCB, which can result in a relatively expensive replacement due to many other components on the PCB. The further analysis process may require destroying the host PCB and performing further physical preparations before the SSD can be powered off, which may enable the manufacturer or customer to reuse the host PCB.
[0020] In some embodiments described herein, a memory device may receive an indication from a host system that the memory device is associated with a fault. The fault may be associated with the host processor's inability to read data from or write data to the memory device. The fault may be associated with the startup of the memory device, where the host processor may not be able to interrogate the memory device based on the fault. The memory device and the host system may be associated with a vehicle. The memory device may receive from the host system a request to initialize a recovery boot mode in response to the fault. The recovery boot mode may provide one or more registers and one or more interfaces that can be used by the memory device and the host processor to assist in the recovery of the memory device. The recovery boot mode may be based on the vendor-specific PCIe recovery boot capabilities of the memory device. The memory device may initialize the recovery boot mode based on the request. The memory device may use backup firmware (e.g., recovery firmware) to initialize the recovery boot mode. The memory device may transmit to the host system status information obtained while operating in the recovery boot mode.
[0021] In some embodiments, by enabling the host system to restart the memory device with backup firmware, the host system may be able to resolve a fault associated with the memory device. For example, the host system may be able to continue communicating with the memory device after restarting the memory device with backup firmware, thereby preventing the host system and / or the memory device from unnecessarily wasting power due to the memory device not responding. By instructing the memory device to initialize a boot recovery mode, the memory device may be able to resolve the fault and continue normal operation without wasting computing resources and / or power, which may improve the overall performance of the host system and the memory device.
[0022] In some embodiments, the recovery boot mode can generate debug information. Even when the vehicle still needs to access the service center to recover the memory device, the debug information can be extracted from the memory device and used to recover the memory device. The use of debug information can eliminate the need to remove the PCB with the memory device from the vehicle and destroy the entire PCB (e.g., destroy the memory device). In the case of removing the PCB from the vehicle, the PCB may not need to be physically destroyed, and the data can be pulled from the memory device and sent for further analysis. The use of debug information can reduce the triage response time because the data is immediately available and can have less latency attributed to the physical preparation of the PCB with the memory device.
[0023] In some embodiments, the ability of the memory device to perform a boot recovery mode in the field can allow the memory device to be recovered and triaged, and reduce the occurrence of having to physically send the memory device to the service center for analysis. When the memory device is sent to the service center, this method may typically require physically destroying the PCB (e.g., the customer application PCB) due to the BGA form factor. The ability to perform a boot recovery mode in the field can allow valuable debug information to be retrieved in a system with a PCB. The ability to perform a startup recovery mode in the field can allow the end customer's application to respond better, and the end customer's application can react to the backup firmware information and better inform the customer to handle certain situations in the event of a vehicle failure. Otherwise, when the memory device in the vehicle becomes unresponsive in the field, the entire subsystem can fail and the vehicle may need to be towed to the service center. Additionally, the ability to recover the memory device using a backup firmware image can reduce the specialized hardware and software typically required to recover an unresponsive memory device during development and customer qualification processes.
[0024] Figure 1 FIG. is a diagram illustrating an example system 100 capable of implementing a recovery mode of a memory device. System 100 may include one or more devices, apparatuses, and / or components for performing the operations described herein. For example, system 100 may include a host system 105 and a memory system 110. Memory system 110 may include a memory system controller 115 and one or more memory devices 120, shown as memory devices 120-1 through 120-N (where N≥1). The memory device may include a local controller 125 and one or more memory arrays 130. Host system 105 may communicate with memory system 110 (e.g., memory system controller 115 of memory system 110) via a host interface 140. Memory system controller 115 and memory device 120 may communicate via respective memory interfaces 145, shown as memory interfaces 145-1 through 145-N (where N≥1).
[0025] System 100 can be any electronic device configured to store data in a memory. For example, System 100 can be a computer, a mobile phone, a wired or wireless communication device, a network device, a server, a device in a data center, a device in a cloud computing environment, a vehicle (e.g., an automobile or an airplane), and / or an Internet of Things (IoT) device. Host system 105 can include host processor 150. Host processor 150 can include one or more processors configured to execute instructions and store data in memory system 110. For example, host processor 150 can include a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or another type of processing component.
[0026] Memory system 110 can be any electronic device or apparatus configured to store data in a memory. For example, memory system 110 can be a hard disk drive, an SSD, a flash memory system (e.g., a NAND flash memory system or a NOR flash memory system), a universal serial bus (USB) drive, a memory card (e.g., a secure digital (SD) card), an auxiliary storage device, a non-volatile memory express (NVMe) device, an embedded multimedia card (eMMC) device, a dual in-line memory module (DIMM), and / or a random access memory (RAM) device, such as a dynamic RAM (DRAM) device or a static RAM (SRAM) device.
[0027] Memory system controller 115 can be any device configured to control the operation of memory system 110 and / or the operation of memory device 120. For example, memory system controller 115 can include control logic, a memory controller, a system controller, an ASIC, an FPGA, a processor, a microcontroller, and / or one or more processing components. In some embodiments, memory system controller 115 can communicate with host system 105 and can direct one or more memory devices 120 to perform memory operations to be executed by that one or more memory devices 120 based on one or more instructions from host system 105. For example, memory system controller 115 can provide instructions to local controller 125 regarding memory operations to be performed by local controller 125 in conjunction with corresponding memory device 120.
[0028] Memory device 120 can include local controller 125 and one or more memory arrays 130. In some embodiments, memory device 120 includes a single memory array 130. In some embodiments, each memory device 120 of memory system 110 can be implemented in a separate semiconductor package or on a separate die that includes that memory device 120, its corresponding local controller 125, and its corresponding memory array 130. Memory system 110 can include multiple memory devices 120.
[0029] The local controller 125 can be any device configured to control the memory operations of the memory device 120 that contains the local controller 125 (e.g., and not the memory operations of other memory devices 120). For example, the local controller 125 can include control logic, a memory controller, a system controller, an ASIC, an FPGA, a processor, a microcontroller, and / or one or more processing components. In some embodiments, the local controller 125 can communicate with the memory system controller 115 and can control operations performed on the memory array 130 coupled to the local controller 125 based on one or more instructions from the memory system controller 115. As an example, the memory system controller 115 can be an SSD controller, and the local controller 125 can be a NAND controller.
[0030] The memory array 130 can include an array of memory cells configured to store data. For example, the memory array 130 can include a non-volatile memory array (e.g., a NAND memory array or a NOR memory array) or a volatile memory array (e.g., a SRAM array or a DRAM array). In some embodiments, the memory system 110 can include one or more volatile memory arrays 135. The volatile memory arrays 135 can include SRAM arrays and / or DRAM arrays and other examples. The one or more volatile memory arrays 135 can be included in the memory system controller 115, included in one or more memory devices 120, and / or included in both the memory system controller 115 and one or more memory devices 120. In some embodiments, the memory system 110 can include both non-volatile memory capable of maintaining stored data after the memory system 110 is powered off and volatile memory (e.g., the volatile memory arrays 135) that requires power to maintain stored data and loses stored data after the memory system 110 is powered off. For example, the volatile memory arrays 135 can cache data read from or to be written to non-volatile memory, and / or can cache instructions to be executed by the controller of the memory system 110.
[0031] The host interface 140 enables communication between the host system 105 (e.g., the host processor 150) and the memory system 110 (e.g., the memory system controller 115). The host interface 140 can include, for example, a Small Computer System Interface (SCSI), a Serial Attached SCSI (SAS), a Serial Advanced Technology Attachment (SATA) interface, a PCIe interface, an NVMe interface, a USB interface, a Universal Flash Storage (UFS) interface, an eMMC interface, a Double Data Rate (DDR) interface, and / or a DIMM interface.
[0032] The memory interface 145 implements communication between the memory system 110 and the memory device 120. The memory interface 145 may include a non-volatile memory interface (e.g., for communicating with non-volatile memory), such as a NAND interface or a NOR interface. Additionally or alternatively, the memory interface 145 may include a volatile memory interface (e.g., for communicating with volatile memory), such as a DDR interface.
[0033] Although the example memory system 110 described above includes a memory system controller 115, in some embodiments, the memory system 110 does not include a memory system controller 115. For example, an external controller (e.g., included in the host system 105) and / or one or more local controllers 125 included in one or more corresponding memory devices 120 may perform the operations described herein as being performed by the memory system controller 115. Additionally, as used herein, "controller" may refer to the memory system controller 115, the local controller 125, or the external controller. In some embodiments, a set of operations described herein as being performed by a controller may be performed by a single controller. For example, the entire set of operations may be performed by a single memory system controller 115, a single local controller 125, or a single external controller. Alternatively, a set of operations described herein as being performed by a controller may be performed by more than one controller. For example, a first subgroup of operations may be performed by the memory system controller 115, and a second subgroup of operations may be performed by the local controller 125. Additionally, depending on the context, the term "memory device" may refer to the memory system 110 or the memory device 120.
[0034] A controller (e.g., the memory system controller 115, the local controller 125, or an external controller) can control operations performed on a memory (e.g., the memory array 130) by, for example, executing one or more instructions. For example, the memory system 110 and / or the memory device 120 can store one or more instructions as firmware in the memory, and the controller can execute that one or more instructions. Additionally or alternatively, the controller can receive one or more instructions from the host system 105 and / or from the memory system controller 115, and can execute that one or more instructions. In some embodiments, a non-transitory computer-readable medium (e.g., volatile memory and / or non-volatile memory) can store a set of instructions (e.g., one or more instructions or code) for the controller to execute. The controller can execute the set of instructions to perform one or more operations or methods described herein. In some embodiments, the controller's execution of the set of instructions causes the controller, the memory system 110, and / or the memory device 120 to perform one or more operations or methods described herein. In some embodiments, hardwired circuitry is used instead of or in combination with one or more instructions to perform one or more operations or methods described herein. Additionally or alternatively, the controller can be configured to perform one or more operations or methods described herein. Instructions are sometimes referred to as "commands".
[0035] For example, a controller (e.g., the memory system controller 115, the local controller 125, or an external controller) can transmit signals to and / or receive signals from a memory (e.g., one or more memory arrays 130) based on one or more instructions, such as to transfer data to (e.g., write or program) all or a portion of the memory (e.g., one or more memory cells, pages, sub-blocks, blocks, or planes of the memory), transfer data from all or a portion of the memory (e.g., read), erase, and / or refresh all or a portion of the memory. Additionally or alternatively, the controller can be configured to control access to the memory and / or provide a translation layer between the host system 105 and the memory (e.g., for mapping logical addresses to physical addresses of the memory array 130). In some embodiments, the controller can translate host interface commands (e.g., commands received from the host system 105) into memory interface commands (e.g., commands for performing operations on the memory array 130).
[0036] In some embodiments, Figure 1 one or more systems, devices, apparatuses, components, and / or controllers can be configured to receive an indication from the host system 105 that the memory device 120 is associated with a failure; receive a request from the host system 105 to initialize a recovery boot mode in response to the failure; initialize the recovery boot mode based on the request; and transmit status information obtained while operating in the recovery boot mode to the host system 105.
[0037] In some embodiments, Figure 1 one or more systems, devices, apparatuses, components, and / or controllers of Figure 1 may be configured to receive a request to initiate a recovery of the memory device 120, where the request is based on a failure associated with the memory device 120; based on the request, initiate a reboot of the memory device 120; and transmit status information obtained when rebooting the memory device 120, where the status information includes information associated with the current operating state of the memory device 120.
[0038] Figure 1 The number and arrangement of the components shown in Figure 1 are provided as an example. In fact, compared to what is shown in Figure 1 , there may be additional components, fewer components, different components, or components arranged in a different manner. Additionally, Figure 1 compared to what is shown in Figure 1 , there may be additional components, fewer components, different components, or components arranged in a different manner. Additionally, Figure 1 two or more of the components shown in Figure 1 may be implemented within a single component, or Figure 1 a single component shown in Figure 1 may be implemented as multiple distributed components. Additionally or alternatively, Figure 1 a set of components (e.g., one or more components) shown in Figure 1 may perform one or more operations described as being performed by another set of components shown in Figure 1 . Figure 1 Figure 1
[0039] Figure 2 FIG. Figure 2 is a diagram illustrating an example 200 of implementing a recovery mode of a memory device. Example 200 may include a host system 105 and a memory device 120. The host system 105 and the memory device 120 may be associated with a vehicle. For example, the memory device 120 may be an automotive SSD.
[0040] As shown by reference numeral 202, the host system 105 may detect a failure associated with the memory device 120. The failure may be associated with the host system 105 being unable to read data from and / or write data to the memory device 120. The failure may be associated with the startup of the memory device 120, where the host system may be unable to interrogate the memory device 120 based on the failure. The host system 105 may detect some errors when attempting to perform an operation with the memory device 120, which may be an indication of the failure.
[0041] As an example, the SSD may fail to boot in the host system 105, which may occur when the SSD has a successful PCIe link but the NVM subsystem initialization fails. The inability to boot in the host system 105 may lead to unpredictable host behavior because the host system may attempt to query the SSD but fail. As another example, an embedded or automotive SSD in a BGA form factor may encounter a failure where the SSD firmware is no longer able to communicate with the host system 105. A critical automotive SSD failure may be attributed to the loss of communication with the host system 105 (e.g., the SSD may no longer function or may otherwise not respond). In these cases, the host system 105 may be able to detect a failure associated with the SSD.
[0042] As shown by component symbol 204, the memory device 120 may receive an indication from the host system 105 that the memory device 120 is associated with a failure. The indication may indicate the type of failure. For example, the indication may indicate that the failure is associated with the host system 105's inability to read data from and / or write data to the memory device 120. As another example, the indication may indicate that the failure is associated with the startup of the memory device 120. In some cases, the indication may not explicitly indicate the type of failure.
[0043] As shown by component symbol 206, the memory device 120 may receive a request from the host system 105 to initialize a recovery of the memory device 120 in response to the failure. For example, the request may be to initialize a recovery boot mode in response to the failure. The host system 105 may request that the memory device 120 initialize a recovery boot mode to resolve the failure. The recovery boot mode may enable the memory device 120 to resolve the failure on-site. The recovery boot mode may be based on the memory device 120's vendor-specific PCIe recovery boot capabilities. In an alternative configuration, the host system 105 may not transmit an indication of the failure to the memory device 120. Instead, the host system 105 may only transmit a request to initialize the recovery boot mode, where the request may be based on a failure detected by the host system 105.
[0044] As shown by component symbol 208, the memory device 120 may initialize a reboot of the memory device 120 based on a request. The memory device 120 may initialize a recovery boot mode based on the request. The recovery boot mode may provide one or more registers and / or one or more interfaces that may be used by the memory device 120 and / or the host system 105 to assist in the recovery of the memory device 120. The reboot may be associated with one or more registers and / or one or more interfaces. The memory device 120 may perform a power cycle of the memory device 120 based on the request. After the power cycle, the memory device 120 may use backup firmware stored on the memory device 120 to initialize the recovery boot mode (or reboot). The memory device 120 may receive debug firmware from the host system 105 and while operating in the recovery boot mode (or when rebooting the memory device 120). The memory device 120 may use the debug firmware to perform debug operations while operating in the recovery boot mode (or when rebooting the memory device 120). The debug operations may generate debug information that may be used to troubleshoot faults. The debug information may indicate various error codes that may be useful in identifying and resolving faults.
[0045] As shown by component symbol 210, the memory device 120 may transmit status information obtained when rebooting the memory device 120 to the host system 105. The status information may include information associated with the current operating state of the memory device 120. The memory device 120 may transmit the status information while operating in the recovery boot mode. The host system 105 may perform some actions based on the status information. For example, when the status information indicates that the recovery boot mode was successful, the host system 105 may continue normal read and write operations with the memory device 120. However, when the status information indicates that the recovery boot mode was not successful, the host system 105 may wait to continue normal read-write operations with the memory device 120.
[0046] In some embodiments, the host system 105 may indicate to the memory device 120 that the host system 105 has detected a fault in the memory device 120. The host system 105 may then request that the memory device 120 initialize into a recovery boot mode using vendor - specific PCIe recovery boot capabilities. The vendor - specific PCIe recovery boot capabilities may provide one or more registers and / or one or more interfaces that the host system 105 and the memory device 120 can use to assist in the recovery boot. The host system 105 may store these capabilities such that the memory device 120 can always boot into the recovery mode until the host system 105 clears the recovery boot mode. The memory device 120 may use backup firmware to support the recovery boot mode. During the recovery boot mode, the memory device 120 may perform only minimal operations, which can ensure that a recovery startup is possible. For example, the memory device 120 may avoid building mapping tables, and the memory device 120 may limit NAND operations, which can allow the memory device 120 to enter and execute the recovery boot mode. During the boot recovery mode, the memory device 120 may download debug firmware in an authenticated manner, which can assist in debug operations.
[0047] In some embodiments, the memory device 120 itself may have the ability to detect faults. For example, the memory device 120 may detect when communication with the host system 105 no longer occurs. In such a case, the memory device 120 itself may initialize the recovery boot mode. The memory device 120 may power cycle and start the recovery firmware. When the recovery is successful, the memory device 120 may continue normal operation with the host system 105.
[0048] In some embodiments, by enabling the host system 105 to restart the memory device 120 with backup firmware, the host system 105 may be able to resolve faults associated with the memory device 120. For example, the host system 105 may be able to continue communication with the memory device 120 after restarting the memory device 120 with backup firmware, thereby preventing the host system 105 and / or the memory device 120 from unnecessarily wasting power due to the memory device 120 not responding. By instructing the memory device 120 to initialize the boot recovery mode, the memory device 120 may be able to resolve the fault and continue normal operation without wasting computing resources and / or power, which can improve the overall performance of the host system 105 and the memory device 120.
[0049] In some embodiments, when the recovery boot mode is unsuccessful in the field and the vehicle needs to be taken to a service center to recover the memory device 120, debug information can be extracted from the memory device 120 and used to recover the memory device 120. The use of debug information can eliminate the need to remove the PCB with the memory device 120 from the vehicle and destroy the entire PCB (e.g., destroy the memory device 120). In the case of removing the PCB from the vehicle, the PCB may not need to be physically destroyed, and data can be pulled from the memory device 120 and sent for further analysis. The use of debug information can reduce the triage response time because the data is immediately available and there are fewer delays attributed to the physical preparation of the PCB with the memory device 120. The ability of the memory device 120 to perform a boot recovery mode in the field can allow the memory device 120 to be recovered and triaged, and reduce the occurrence of having to physically take the memory device 120 to a service center for analysis. The ability to perform a boot recovery mode in the field can allow valuable debug information to be retrieved in a system with a PCB, which can avoid having to physically destroy the PCB.
[0050] As indicated above, Figure 2 is provided as an example. Other examples may differ from those Figure 2 described with respect to
[0051] Figure 3 is a diagram illustrating Example 300 of implementing a recovery mode of a memory device.
[0052] As shown by component symbol 302, a host system (e.g., host system 105) can check whether an SSD (e.g., memory device 120) is accessible and operating correctly. For example, the host system can check whether the host system is able to communicate with the SSD. The SSD can be an embedded automotive SSD. As shown by component symbol 304, when the SSD is accessible and operating correctly, the host system and the SSD can continue normal operation. As shown by component symbol 306, when the SSD is not accessible and / or not operating correctly, the host system can request the SSD to boot into a recovery mode using vendor-specific capabilities. The vendor-specific capabilities can be recovery-mode vendor-specific PCIe capabilities that can allow the SSD to start up in the recovery mode. As shown by component symbol 308, the SSD can perform a power cycle based on a request received from the host system. As shown by component symbol 310, after power-on, the SSD can identify the recovery mode and use recovery firmware to boot into the recovery mode. The SSD can use recovery firmware instead of normal firmware. Compared with normal firmware, recovery firmware may have limited functionality. For example, recovery firmware can avoid building tables and / or may only support limited NAND operations, which can allow the SSD to enter the recovery mode and operate in the recovery mode. The recovery firmware can be executed after a power cycle of the SSD. As shown by component symbol 312, during the recovery mode, the SSD firmware can start in a debug / recovery mode to provide critical data status to the host system. As shown by component symbol 314, the host system can act based on the SSD debug data and status information. The host system can perform various actions based on the SSD debug data and status information, which can involve providing the SSD debug data and status information for display on a user interface. After the SSD recovers and continues normal operation, the host system can continue to monitor the SSD.
[0053] In some embodiments, the host system can restart the SSD with recovery firmware on-site. The recovery firmware can enable the SSD to attempt an on-site recovery, which can be useful when the SSD is an embedded automotive SSD. In this case, the SSD can be integrated with the vehicle, and when the SSD does not respond, the vehicle may not be usable. The host system can attempt to recover the SSD on-site instead of automatically requiring the vehicle itself to be taken to a service center. When the on-site recovery is unsuccessful, the vehicle can be taken to a service center.
[0054] As indicated above, Figure 3 is provided as an example. Other examples may be different from those Figure 3 described.
[0055] Figure 4is a flowchart of an example method 400 associated with implementing a recovery mode of a memory device. In some embodiments, a memory device (e.g., memory device 120) may execute or may be configured to execute method 400. In some embodiments, another device or group of devices (e.g., system 100) independent of or including the memory device may execute or may be configured to execute method 400. Additionally or alternatively, one or more components of the memory device (e.g., controller 125) may execute or may be configured to execute method 400. Thus, the means for executing method 400 may include the memory device and / or one or more components of the memory device. Additionally or alternatively, a non-transitory computer-readable medium may store one or more instructions that, when executed by a memory device (e.g., controller 125 of memory 120), cause the memory device to execute method 400.
[0056] As Figure 4 shown, method 400 may include receiving, by the memory device from a host system, a request to initialize recovery of the memory device, where the request is based on a fault associated with the memory device (block 410). As Figure 4 further shown, method 400 may include initializing a reboot of the memory device based on the request (block 420). As Figure 4 further shown, method 400 may include transmitting status information obtained when rebooting the memory device, where the status information includes information associated with a current operating state of the memory device (block 430).
[0057] Method 400 may include additional aspects, such as any single aspect or any combination of aspects described below and / or in combination with one or more other methods or operations described elsewhere herein.
[0058] In a first aspect, a recovery boot mode provides one or more registers and one or more interfaces that may be used by the memory device and the host system to assist in the recovery of the memory device.
[0059] In a second aspect, either alone or in combination with the first aspect, method 400 includes performing a power cycle of the memory device based on the request and, after the power cycle, initializing the recovery boot mode using backup firmware stored on the memory device.
[0060] In a third aspect, either alone or in combination with one or more of the first and second aspects, the fault is associated with the host system's inability to read data from or write data to the memory device.
[0061] In a fourth aspect, either alone or in combination with one or more of the first through third aspects, the fault is associated with the startup of the memory device and the host system is unable to interrogate the memory device based on the fault.
[0062] In a fifth aspect, either alone or in combination with one or more of the first to fourth aspects, method 400 includes receiving debug firmware from a host system and while operating in a recovery boot mode, and performing a debug operation using the debug firmware while operating in the recovery boot mode.
[0063] In a sixth aspect, either alone or in combination with one or more of the first to fifth aspects, the recovery boot mode is based on a vendor-specific PCIe recovery boot capability of the memory device.
[0064] Although Figure 4 illustrative boxes of method 400 are shown, in some embodiments, method 400 may include additional boxes, fewer boxes, different boxes, or boxes arranged in a different manner compared to those depicted in Figure 4 . Additionally or alternatively, two or more boxes of method 400 may be performed in parallel. Method 400 is an example of a method executable by one or more of the devices described herein. These one or more devices may execute or may be configured to execute one or more other methods based on the operations described herein.
[0065] In some embodiments, a memory device includes one or more components configured to: receive an indication from a host system that the memory device is associated with a fault; receive from the host system a request to initialize a recovery boot mode in response to the fault; initialize the recovery boot mode based on the request; and transmit status information obtained while operating in the recovery boot mode to the host system.
[0066] In some embodiments, a method includes: a memory device receiving from a host system a request to initialize a recovery of the memory device, where the request is based on a fault associated with the memory device; initializing, based on the request, a reboot of the memory device; and transmitting status information obtained while rebooting the memory device, where the status information includes information associated with a current operating state of the memory device.
[0067] In some embodiments, a system includes a host system configured to: detect a fault associated with a memory device and transmit to the memory device a request to initialize a recovery of the memory device in response to the fault; and the memory device configured to: receive the request from the host system, initialize a reboot of the memory device based on the request, and transmit status information obtained while rebooting the memory device, where the status information includes information associated with a current operating state of the memory device.
[0068] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations may be made in light of the above disclosure, or may be acquired from the practice of the embodiments described herein.
[0069] As used herein, "meeting a threshold" may, depending on the context, refer to a value that is greater than a threshold, greater than or equal to a threshold, less than a threshold, less than or equal to a threshold, equal to a threshold, not equal to a threshold, or the like.
[0070] Even if specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of the embodiments described herein. Many of these features may be combined in ways not specifically recited in the claims and / or not disclosed in the specification. For example, the present disclosure encompasses each dependent claim in a claim set, as well as each other individual claim in the claim set, and each combination of multiple claims in the claim set. As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including a single member. For example, "at least one of a, b, or c" is intended to cover a, b, c, a + b, a + c, b + c, and a + b + c, as well as any combination with multiples of the same element (e.g., a + a, a + a + a, a + a + b, a + a + c, a + b + b, a + c + c, b + b, b + b + b, b + b + c, c + c, and c + c + c, or any other ordering of a, b, and c).
[0071] When a "component" or "one or more components" (or another element, such as a "controller" or "one or more controllers") is described or claimed as performing multiple operations or being configured to perform multiple operations (within a single claim or across multiple claims), this language is intended to broadly cover a variety of architectures and environments. For example, unless otherwise explicitly stated (e.g., by using "a first component" and "a second component" or other language that differentiates components in the claims), this language is intended to cover a single component that performs or is configured to perform all operations, a group of components that jointly perform or are configured to perform all operations, a first component that performs or is configured to perform a first operation and a second component that performs or is configured to perform a second operation, or any combination of components that perform or are configured to perform the operations. For example, when a claim has the form "one or more components configured to: perform X; perform Y; and perform Z", the claim should be interpreted to mean "one or more components configured to perform X; one or more (possibly different) components configured to perform Y; and one or more (also possibly different) components configured to perform Z".
[0072] Unless explicitly described, any element, act, or instruction used herein should not be construed as critical or essential. Also, as used herein, the article "a / an" is intended to include one or more items and may be used interchangeably with "one or more." Additionally, as used herein, the article "the" is intended to include one or more items referred to in conjunction with the article "the" and may be used interchangeably with "the one or more." If only one item is desired, then the phrases "only one," "single," or similar language is used. Further, as used herein, the term "has / have / having" or the like is intended to be an open-ended term that does not limit the element it modifies (e.g., an element that "has" A may also have B). Additionally, the phrase "based on" is intended to mean "at least partially based on" unless otherwise explicitly stated. As used herein, the term "plurality" may be replaced with "a plurality of," and vice versa. Also, as used herein, the term "or" is intended to be inclusive when used in a series and may be used interchangeably with "and / or" unless otherwise explicitly stated (e.g., if used in combination with "either" or "only one of...").
Claims
1. A memory device, comprising: One or more components configured to: Receive an indication from a host system that the memory device is associated with a fault; Receive from the host system a request to initialize a recovery of the memory device in response to the fault; Initialize a reboot of the memory device based on the request; And Transmit status information obtained when rebooting the memory device, where the status information includes information associated with the current operating state of the memory device.
2. The memory device according to claim 1, wherein the reboot is associated with one or more registers and one or more interfaces that can be used by the memory device and the host system to assist in the recovery of the memory device.
3. The memory device according to claim 1, wherein the one or more components are further configured to: Perform a power cycle of the memory device based on the request; and After the power cycle, initialize the reboot using backup firmware stored on the memory device.
4. The memory device according to claim 1, wherein the fault is associated with the host system being unable to read data from or write data to the memory device.
5. The memory device according to claim 1, wherein the fault is associated with the startup of the memory device, and the host system is unable to interrogate the memory device based on the fault.
6. The memory device according to claim 1, wherein the one or more components are further configured to: Receive debug firmware from the host system and when rebooting the memory device; and When rebooting the memory device, perform debug operations using the debug firmware.
7. The memory device according to claim 1, wherein the reboot is based on the vendor - specific Peripheral Component Interconnect Express (PCIe) recovery boot capability of the memory device.
8. The memory device according to claim 1, wherein the memory device is associated with a vehicle.
9. A method, comprising: Receiving, by a memory device from a host system, a request to initialize a recovery of the memory device, where the request is based on a fault associated with the memory device; Initializing a reboot of the memory device based on the request; and Transmitting status information obtained when rebooting the memory device, where the status information includes information associated with the current operating state of the memory device.
10. The method according to claim 9, wherein the reboot is associated with one or more registers and one or more interfaces that can be used by the memory device and the host system to assist in the recovery of the memory device.
11. The method according to claim 9, further comprising: Performing a power cycle of the memory device based on the request; And After the power cycle, initializing the reboot using backup firmware stored on the memory device.
12. The method according to claim 9, wherein the failure is associated with the host system being unable to read data from or write data to the memory device.
13. The method according to claim 9, wherein the failure is associated with the startup of the memory device, and the host system is unable to interrogate the memory device based on the failure.
14. The method according to claim 9, further comprising: Receiving debug firmware from the host system and when rebooting the memory device; And Performing a debug operation using the debug firmware when rebooting the memory device.
15. The method according to claim 9, wherein the reboot is based on the vendor-specific Peripheral Component Interconnect Express (PCIe) recovery boot capability of the memory device.
16. A system, comprising: A host system configured to: Detect a failure associated with a memory device, and Transmit a request to the memory device to initiate a recovery of the memory device in response to the failure; and The memory device configured to: Receive the request from the host system, Initiate a reboot of the memory device based on the request, and Transmit status information obtained when rebooting the memory device, wherein the status information includes information associated with the current operating state of the memory device.
17. The system according to claim 16, wherein the reboot is associated with one or more registers and one or more interfaces that can be used by the memory device and the host system to assist in the recovery of the memory device, and the reboot is based on the vendor-specific Peripheral Component Interconnect Express (PCIe) recovery boot capability of the memory device.
18. The system according to claim 16, wherein the memory device is further configured to: Perform a power cycle of the memory device based on the request; and After the power cycle, initialize the recovery boot mode using backup firmware stored on the memory device.
19. The system according to claim 16, wherein: The failure is associated with the host system being unable to read data from or write data to the memory device; or The failure is associated with the startup of the memory device, and the host system is unable to interrogate the memory device based on the failure.
20. The system according to claim 16, wherein the memory device is further configured to: Receive debug firmware from the host system and when rebooting the memory device; and Perform a debug operation using the debug firmware when rebooting the memory device.