Memory subsystem recovery in response to host hardware recovery signals
By enabling backup firmware using the hardware recovery signal of the host system assertion in the memory subsystem, the difficulty of debugging and recovery of the memory subsystem in the event of failure is solved, and a non-destructive and efficient debugging and recovery process is achieved.
Patent Information
- Application Number
- CN202411651654.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-25
- Filing Date
- 2024-11-19
- Publication Date
- 2025-05-20
AI Technical Summary
Existing memory subsystems are difficult to effectively debug and restore when encountering failures or errors, especially when the memory subsystem cannot communicate with the host system, debugging and recovery becomes difficult and destructive.
Enable backup firmware to restore unresponsive memory subsystems by using the hardware recovery signal of the host system assertion. The backup firmware is loaded after the power cycle event, providing debugging functionality, allowing the memory subsystem to communicate status information to the host system and perform other debugging operations.
It realizes debugging and recovery while the memory subsystem is kept in place, avoids the need for physical destruction of packaging, reduces dependence on dedicated hardware and software, and retains physical space and resources.
Smart Images

Figure CN120020737A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to memory subsystems, and more particularly, to recovering an unresponsive memory subsystem using backup firmware enabled by a host-initiated hardware recovery signal. Background Art
[0002] A memory subsystem may include one or more memory devices that store data. The memory devices may be, for example, non-volatile memory devices and volatile memory devices. Generally, a host system may utilize the memory subsystem to store data at and retrieve data from the memory devices. Summary of the Invention
[0003] In one aspect, the present disclosure provides a memory subsystem including: a memory device; a processing device operably coupled to the memory device to perform operations including: executing a restart sequence in response to the occurrence of a power cycle event; determining whether a recovery signal is asserted by a host system via a sideband interface; in response to determining that the recovery signal is asserted, retrieving and loading a backup firmware image of the memory subsystem; and booting the memory subsystem in a debug operation mode using the backup firmware image.
[0004] In another aspect, the present disclosure further provides a method including: executing a restart sequence of a memory subsystem in response to the occurrence of a power cycle event; determining whether a recovery signal is asserted by a host system via a sideband interface; in response to determining that the recovery signal is asserted, retrieving and loading a backup firmware image of the memory subsystem; and booting the memory subsystem in a debug operation mode using the backup firmware image.
[0005] In another aspect, the present disclosure further provides a non-transitory computer-readable storage medium including instructions that, when executed by a processing device, cause the processing device to perform operations including: executing a restart sequence of a memory subsystem in response to the occurrence of a power cycle event; determining whether a recovery signal is asserted by a host system via a sideband interface; in response to determining that the recovery signal is asserted, retrieving and loading a backup firmware image of the memory subsystem; and booting the memory subsystem in a debug operation mode using the backup firmware image. Brief Description of the Drawings
[0006] The present disclosure will be more fully understood from the detailed description given below and from the accompanying drawings of various embodiments of the present disclosure.
[0007] Figure 1 Illustrate an example computing system including a memory subsystem according to some embodiments of the present disclosure.
[0008] Figure 2is a block diagram of a computing system that illustrates recovery of an unresponsive memory subsystem using backup firmware enabled by a hardware recovery signal initiated by a host, in accordance with some embodiments of the present disclosure.
[0009] Figure 3 is a sequence diagram that illustrates recovery of an unresponsive memory subsystem using backup firmware enabled by a hardware recovery signal initiated by a host, in accordance with some embodiments of the present disclosure.
[0010] Figure 4 is a flow diagram of an example method for recovery of an unresponsive memory subsystem using backup firmware enabled by a hardware recovery signal initiated by a host, in accordance with some embodiments of the present disclosure.
[0011] Figure 5 is a block diagram of an example computer system in which embodiments of the present disclosure may operate. DETAILED DESCRIPTION
[0012] Aspects of the present disclosure relate to recovery of an unresponsive memory subsystem using backup firmware enabled by a hardware recovery signal initiated by a host. The memory subsystem may be a storage device, a memory module, or a combination of a storage device and a memory module. Example storage devices and memory modules are described below in connection with Figure 1 Generally, a host system may utilize a memory subsystem that includes one or more components, such as a memory device that stores data. The host system may provide data stored at the memory subsystem and may request data retrieved from the memory subsystem.
[0013] The memory subsystem may include high density non-volatile memory devices where it is desirable to retain data when there is no power supply to the memory device. For example, NAND memory (such as 3D flash NAND memory) provides storage in a compact high density configuration. A non-volatile memory device is an encapsulation of one or more dies, each die including one or more planes. For some types of non-volatile memory devices (such as NAND memory), each plane includes a set of physical blocks. Each block includes a set of pages. Each page includes a set of memory cells (“cells”). A cell is an electronic circuit that stores information. Depending on the cell type, a cell may store one or more binary information bits and have various logical states associated with the number of bits stored. The logical states may be represented by binary values such as “0” and “1” or combinations of such values.
[0014] An example of a memory subsystem is a solid state drive (SSD) that includes one or more non-volatile memory devices and a memory subsystem controller for managing the non-volatile memory devices. The memory devices may be composed of bits arranged in a two-dimensional or three-dimensional grid. Memory cells are formed on a silicon wafer in an array of columns (also referred to hereinafter as bit lines) and rows (also referred to hereinafter as word lines). A word line may refer to one or more rows of memory cells of the memory device, which are used in conjunction with one or more bit lines to generate an address for each of the memory cells. The intersection of a bit line and a word line constitutes the address of a memory cell. A block refers hereinafter to a unit of the memory device for storing data and may include a group of memory cells, a group of word lines, a word line, or an individual memory cell. One or more blocks may be grouped together to form a separate partition (e.g., a plane) of the memory device to allow concurrent operations to occur on each plane.
[0015] Like any electronic circuit, a memory subsystem is vulnerable to various types of errors or faults that can affect performance and / or operability. For example, a fault in the firmware executed on the memory subsystem can cause the memory subsystem to become unresponsive, an input / output error can prevent communication with the host system, or a fault in the non-volatile memory device of the memory subsystem can impede the storage or retention of data. Debugging is a methodical process of identifying and reducing the number of defects (i.e., "bugs") in the memory subsystem that cause the above errors or faults. Various debugging techniques can be used to detect anomalies, evaluate their impact, and schedule hardware changes, firmware upgrades, or complete updates of the memory subsystem. The goals of debugging include identifying and fixing bugs in the system (e.g., logic or synchronization problems in the firmware or design errors in the hardware) and collecting system state information, such as information about the operation of the memory subsystem, which can then be used to analyze the memory subsystem to find ways to recover from faults, improve its performance, or optimize other important characteristics.
[0016] In some systems, debugging operations or other analysis of a memory subsystem are performed on a separate computing device (e.g., a host computing system) communicatively coupled to the memory subsystem via a communication pipeline. The communication pipeline can be implemented using any of a variety of techniques and can include, for example, a Peripheral Component Interconnect Express (PCIe) bus or some other type of communication mechanism. When performing debugging operations, these conventional systems transfer debugging information (e.g., system status information, statistics, runtime analysis, etc.) from the memory subsystem to the host system via the PCIe bus. To transfer debugging information via the PCIe bus, the data must be formatted according to a specific specification, such as the Non-Volatile Memory Express (NVMe) specification. Additionally, the transfer of debugging information utilizes the bandwidth of the PCIe bus, which then cannot be used for other memory subsystem and host operations. Certain types of faults in the memory subsystem prevent the PCIe bus from being used for any communication, let alone transferring debugging information. In some cases, the memory subsystem can be permanently integrated into some other system, such as by attaching the memory subsystem to a printed circuit board or other substrate containing other electrical components using ball grid array packaging technology. This form factor is common in some embedded systems (e.g., automotive implementations) and can include, for example, an entire media control unit of an automobile. When the memory subsystem is permanently integrated with other components and the PCIe bus is not available for transferring debugging information, the only option for inspection and debugging may be to physically destroy the package in order to access the memory subsystem. This method is expensive, destructive, and should be avoided if possible. However, the inability to perform debugging can lead to the inoperability of the memory subsystem as well as other components integrated with the memory subsystem.
[0017] Aspects of the present disclosure address the above and other deficiencies by implementing a backup firmware recovery for an unresponsive memory subsystem enabled by a host-initiated hardware recovery signal. In response to an unexpected error or other fault, the memory subsystem may become unresponsive, such that the host system can no longer communicate with the memory subsystem. In this example, the host system may assert a hardware recovery signal transmitted to the memory subsystem in a sideband interface (e.g., general-purpose input / output (GPIO) pins physically separate from an in-band interface (e.g., a PCIe bus)). The host system may further initiate a power cycle of the memory subsystem by temporarily removing power applied to the memory subsystem and then restoring power after a period of time. When restarted in response to the power cycle, the memory subsystem may sample the recovery signal at the GPIO pins. Under normal circumstances, when the recovery signal is not asserted by the host system, the memory subsystem may continue to retrieve a default firmware image from an on-chip database, load this default firmware, and continue normal operation. However, if the memory subsystem determines that the recovery signal is asserted by the host system, then the memory subsystem may instead retrieve a backup firmware image from the on-chip database and load this backup firmware. In one embodiment, the backup firmware image enables debug functionality that is not normally available with the default firmware. For example, the backup firmware image may allow the memory subsystem to communicate status information to the host system, unlock certain bi-directional ports in the memory subsystem to allow debug operations to be performed, grant access to certain vendor-specific logs in the memory subsystem that are normally restricted, allow firmware updates to be downloaded to fix identified vulnerabilities, or enable other debug functionality.
[0018] Advantages of the methods described herein include (but are not limited to) improved debugging in the memory subsystem. By leveraging a host-initiated recovery signal to control what firmware is loaded onto the memory subsystem after a power cycle, the host system gains a degree of control over recovery and debugging after encountering an unresponsive memory subsystem. This recovery and debugging can be performed with the memory subsystem remaining in place and without breaking the permanent bond (e.g., BGA) in the package. Additionally, this can reduce the need for dedicated hardware and software to be added to the memory subsystem, which preserves physical space and existing resources for other operations of the memory subsystem.
[0019] Figure 1 An example computing system 100 including a memory subsystem 110 is illustrated in accordance with some embodiments of the present disclosure. The memory subsystem 110 may include media such as one or more volatile memory devices (e.g., memory device 140), one or more non-volatile memory devices (e.g., one or more memory devices 130), or a combination thereof.
[0020] The memory subsystem 110 can be a storage device, a memory module, or a hybrid of a storage device and a memory module. Examples of storage devices include solid state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controllers (eMMCs), universal flash storage (UFS) drives, secure digital (SD) cards, and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small DIMMs (SO-DIMMs), and various types of non-volatile dual in-line memory modules (NVDIMMs).
[0021] The computing system 100 can be a computing device such as a desktop computer, a laptop computer, a network server, a mobile device, a vehicle (such as an airplane, a drone, a train, an automobile, or other transportation vehicle), an Internet of Things (IoT) enabled device, an embedded computer (such as an embedded computer included in a vehicle, an industrial device, or a networked commercial device), or such a computing device that includes a memory and a processing device.
[0022] The computing system 100 can include a host system 120 coupled to one or more memory subsystems 110. In some embodiments, the host system 120 is coupled to different types of memory subsystems 110. Figure 1 An example of a host system 120 coupled to one memory subsystem 110 is illustrated. As used herein, "coupled to" or "coupled with" generally refers to a connection between components, which can be an indirect communication connection or a direct communication connection (e.g., without an intervening component), whether wired or wireless, including connections such as electrical, optical, magnetic, etc.
[0023] The host system 120 can include a processor chipset and a software stack executed by the processor chipset. The processor chipset can include one or more cores, one or more caches, a memory controller (such as an NVDIMM controller), and a storage protocol controller (such as a PCIe controller, a SATA controller). The host system 120 uses the memory subsystem 110 to write data to the memory subsystem 110 and read data from the memory subsystem 110, for example.
[0024] The host system 120 can be coupled to the memory subsystem 110 via a physical host interface 122. Examples of the physical host interface 122 include (but are not limited to) Serial Advanced Technology Attachment (SATA) interface, Peripheral Component Interconnect Express (PCIe) interface, Universal Serial Bus (USB) interface, Fibre Channel, Serial Attached SCSI (SAS), Double Data Rate (DDR) memory bus, Small Computer System Interface (SCSI), Dual In-line Memory Module (DIMM) interface (e.g., DIMM socket interface supporting Double Data Rate (DDR)), etc. The physical host interface 122 can be used to transfer data between the host system 120 and the memory subsystem 110. When the memory subsystem 110 is coupled to the host system 120 via a PCIe interface, the host system 120 can further utilize the Non-Volatile Memory Express (NVMe) interface to access the memory components (e.g., one or more memory devices 130). The physical host interface 122 can provide an interface for passing control, address, data, and other signals between the memory subsystem 110 and the host system 120. Figure 1 The memory subsystem 110 is illustrated as an example. Generally, the host system 120 can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.
[0025] The memory devices 130, 140 can include any combination of different types of non-volatile memory devices and / or volatile memory devices. The volatile memory device (e.g., the memory device 140) can be (but is not limited to) random access memory (RAM), such as dynamic random access memory (DRAM) and synchronous dynamic random access memory (SDRAM).
[0026] Some examples of non-volatile memory devices (e.g., the memory device 130) include NAND-type flash memory and write-in-place memory, such as three-dimensional cross-point (“3D cross-point”) memory. The cross-point array of non-volatile memory can perform bit storage based on the change of bulk resistance combined with a stackable cross-gate format data access array. Additionally, compared with many flash-based memories, the cross-point non-volatile memory can perform write-in-place operations, in which non-volatile memory cells can be programmed without prior erasure of the non-volatile memory cells. NAND-type flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND).
[0027] Each of the memory devices 130 may include one or more memory cell arrays. A type of memory cell, such as a single-level cell (SLC), may store one bit per cell. Other types of memory cells, such as multi-level cells (MLC), triple-level cells (TLC), and quad-level cells (QLC), may store multiple bits per cell. In some embodiments, each of the memory devices 130 may include one or more memory cell arrays, such as SLC, MLC, TLC, QLC, or any combination thereof. In some embodiments, a particular memory device may include an SLC portion and an MLC portion, a TLC portion, or a QLC portion of memory cells. The memory cells of the memory devices 130 may be grouped into pages, which may refer to a logical unit of the memory device for storing data. For some types of memory (e.g., NAND), pages may be grouped to form blocks.
[0028] Although non-volatile memory components such as 3D cross-point arrays of non-volatile memory cells and NAND-type flash memories (e.g., 2D NAND, 3D NAND) are described, the memory devices 130 may be based on any other type of non-volatile memory, such as read-only memory (ROM), phase change memory (PCM), self-selecting memory, other chalcogenide-based memories, ferroelectric transistor random access memory (FeTRAM), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), spin transfer torque (STT)-MRAM, conductive bridge RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), nor flash memory, electrically erasable programmable read-only memory (EEPROM).
[0029] The memory subsystem controller 115 (or simply referred to as the controller 115) may communicate with the memory devices 130 to perform operations such as reading data, writing data, or erasing data and other such operations at the memory devices 130. The memory subsystem controller 115 may include hardware, such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The hardware may include digital circuitry having dedicated (i.e., hard-coded) logic for performing the operations described herein. The memory subsystem controller 115 may be a microcontroller, dedicated logic circuitry (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), or other suitable processor.
[0030] The memory subsystem controller 115 may include a processor 117 (e.g., a processing device) configured to execute instructions stored in local memory 119. In the illustrative example, the local memory 119 of the memory subsystem controller 115 includes an embedded memory configured to store instructions for performing various processes, operations, logic flows, and routines for controlling the operation of the memory subsystem 110, including handling communication between the memory subsystem 110 and the host system 120.
[0031] In some embodiments, the local memory 119 may include memory registers for storing memory pointers, fetching data, etc. The local memory 119 may also include a read-only memory (ROM) for storing microcode. Although the example memory subsystem 110 in Figure 1 has been illustrated as including a memory subsystem controller 115, in another embodiment of the present disclosure, the memory subsystem 110 does not include a memory subsystem controller 115 but may instead rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem).
[0032] Generally, the memory subsystem controller 115 may receive commands or operations from the host system 120 and may convert the commands or operations into instructions or appropriate commands to achieve the desired access to the memory device 130. The memory subsystem controller 115 may be responsible for other operations such as wear-leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address translation between logical addresses (e.g., logical block addresses (LBAs), namespaces) and physical addresses (e.g., physical block addresses) associated with the memory device 130. The memory subsystem controller 115 may further include host interface circuitry for communicating with the host system 120 via a physical host interface 122. The host interface circuitry may convert commands received from the host system into command instructions to access the memory device 130 and convert responses associated with the memory device 130 into information for the host system 120.
[0033] The memory subsystem 110 may also include additional circuitry or components not shown. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuitry (e.g., row decoders and column decoders) that may receive addresses from the memory subsystem controller 115 and decode the addresses to access the memory device 130.
[0034] In some embodiments, the memory device 130 includes a local media controller 135 that operates in conjunction with the memory subsystem controller 115 to perform operations on one or more memory cells of the memory device 130. An external controller (such as the memory subsystem controller 115) may manage the memory device 130 externally (e.g., perform media management operations on the memory device 130). In some embodiments, the memory device 130 is a managed memory device, which is a raw memory device (such as the memory array 104) having control logic (such as the local controller 135) for media management within the same memory device package. An example of a managed memory device is a managed NAND (MNAND) device. For example, the memory devices 130 may each represent a single die on which certain control logic (such as the local media controller 135) is embodied. In some embodiments, one or more components of the memory subsystem 110 may be omitted.
[0035] In one embodiment, the memory subsystem 110 includes a recovery manager component 113 that coordinates the use of backup firmware enabled by a host-initiated hardware recovery signal received via a sideband interface 124 to recover an unresponsive memory subsystem 110, where the sideband interface 124 is separate from the primary physical host interface 122. In response to an unexpected error or other failure, the memory subsystem 110 may become unresponsive, such that the host system 120 can no longer communicate with the memory subsystem 110 via the physical host interface 122. In this example, the host system 120 may assert a hardware recovery signal transmitted to the memory subsystem 110 via the sideband interface 124, and the sideband interface 124 may include, for example, general-purpose input / output (GPIO) pins that are physically separate from the PCIe bus used as the primary physical host interface 122. The host system 120 may further initiate a power cycle of the memory subsystem 110 by temporarily removing power applied to the memory subsystem 110 and then restoring power after a period of time. When restarted in response to the power cycle, the recovery manager 113 of the memory subsystem may sample the recovery signal at the sideband interface 124. Under normal circumstances, when the recovery signal is not asserted by the host system 120, the recovery manager 113 may continue to retrieve a default firmware image from an on-chip database (such as local memory 119, memory device 130, or memory device 140), load this default firmware, and continue normal operation. However, if the recovery manager 113 determines that the recovery signal is asserted by the host system 120, then the recovery manager 113 may instead retrieve a backup firmware image from the on-chip database and load this backup firmware. In one embodiment, the backup firmware image enables debug functionality that is not normally available with the default firmware. For example, the backup firmware image may allow the memory subsystem 110 to communicate status information to the host system 120, unlock certain bi-directional ports in the memory subsystem 110 to allow debug operations to be performed, grant access to certain vendor-specific logs in the memory subsystem 110 that are normally restricted, allow firmware updates to be downloaded to fix identified vulnerabilities, or enable other debug functionality. More details regarding the operation of the recovery manager 113 are described below.
[0036] Figure 2It is a block diagram of a computing system that illustrates the use of backup firmware enabled by a hardware recovery signal initiated by a host to recover an unresponsive memory subsystem according to some embodiments of the present disclosure. In one embodiment, the host system 120 is coupled to the memory subsystem 110 via a physical host interface 122 (e.g., a PCIe bus) and via a sideband interface 124. The host system 120 may include corresponding bus ports for each interface and corresponding interface drivers. These interface drivers may provide programming interfaces to control and manage the signals transmitted at the corresponding ports. The host system 120 may further include a host operating system and / or one or more client applications running on top of the host operating system that can access the bus ports via the corresponding drivers (e.g., for sending and receiving data). For example, the host operating system or a client application may call routines in any of the drivers, and the drivers may in turn issue one or several corresponding commands to the associated ports. In one embodiment, the host system 120 communicates with a power controller 225 that supplies power to the memory subsystem 110. After determining that the memory subsystem 110 is unresponsive (i.e., unresponsive to requests or commands sent via the physical host interface 122), the processing logic executed on the host system 120 may initiate a power cycle event by sending a command to the power controller 225 to temporarily stop the power supply to the memory subsystem 110 and then resume the power supply to the memory subsystem 110 after a period of time. Before, during, or after this power cycle event, the host system 120 may assert a recovery signal via the sideband interface 124, and the recovery signal may be detected by the recovery manager 113 when the memory subsystem 110 restarts after the power cycle event.
[0037] In one embodiment, the memory subsystem 110 also includes corresponding bus ports to which the PCIe bus 122 and the sideband interface 124 are coupled. Similar to the host system 120, the bus ports in the memory subsystem 110 may be controlled by corresponding device drivers. As described above, the memory subsystem controller 115 includes a recovery manager 113, which may control sampling the recovery signal at the sideband interface 124, for example, by using the corresponding device drivers and bus ports. In response to determining that the recovery signal is not asserted on the sideband interface 124, the recovery manager 113 may retrieve a default firmware image 212 from the on-chip database 210, load this default firmware, and continue normal operation. The database 210 may include the local memory 119, one of the memory devices 130, or the memory device 140, as described above with respect to Figure 1 the description. However, in response to determining that the recovery signal has been asserted on the sideband interface 124, the recovery manager 113 may instead retrieve a backup firmware image 214 from the database 210 and load this backup firmware. As described herein, the backup firmware image 214 is different from the default firmware image 212 and enables debug functionality that is not normally available with the default firmware.
[0038] Figure 3 This is a sequence diagram illustrating the recovery of an unresponsive memory subsystem using backup firmware enabled by a hardware recovery signal initiated by a host, according to some embodiments of the present disclosure. Sequence diagram 300 illustrates an embodiment of a data exchange procedure executed between a memory subsystem 110 and a host system 120. At operation 302, the host system 120 sends a communication request to the memory subsystem 110. The communication request can be any type of command or request, such as a memory access request (e.g., read command, program command), a status request, a vendor-specific command, etc. In one embodiment, the communication request is sent to the memory subsystem 110 via the PCIe bus 122. Under normal circumstances, the memory subsystem 110 will respond to the communication request (e.g., with requested data, with an acknowledgment message, with an error message). However, in certain situations, such as if the memory subsystem 110 has suffered a certain fault or error that affects performance and / or operability, then the memory subsystem 110 may be unresponsive. Therefore, if the host system 120 does not receive a response after a certain period of time or after a specific number of requests, then at operation 304, the host system 120 may determine that the memory subsystem 110 is unresponsive.
[0039] After determining that the memory subsystem 110 is unresponsive, at operation 306, the host system 120 may assert a recovery signal on the sideband interface 124 and at operation 308, may initiate a power cycle of the memory subsystem 110. In one embodiment, the host system 120 asserts the recovery signal by driving a voltage representing a logic "high" state (e.g., a voltage above a threshold) on the sideband interface 124, which can be received at the memory subsystem 110 (e.g., at the GPIO pin to which the sideband interface 124 is connected). In one embodiment, the host system 120 may initiate a power cycle event by sending a command to the power controller 225 to temporarily stop the power supply to the memory subsystem 110 and then restore the power to the memory subsystem 110 after a period of time.
[0040] At operation 310, the memory subsystem 110 starts after power restoration. In addition to the numerous diagnostic checks to be performed, the recovery manager 113 of the memory subsystem controller 115 may sample a recovery signal as part of the startup process. In response to determining that the recovery signal has been asserted on the sideband interface 124, at operation 312, the recovery manager 113 may retrieve the backup firmware image 214 from the database 210 and load this backup firmware. As described herein, the backup firmware image 214 is different from the default firmware image 212 and enables debug functionality that is not normally available with the default firmware. At operation 314, the recovery manager 113 provides status information to the host system 120. This status information may include an indication that the memory subsystem has restarted and that the backup firmware has been loaded, as well as other information. The status information may include an indication of a fault or error (if known) that has caused the memory subsystem 110 to become unresponsive. Receiving the status information may indicate to the host system 120 that the memory subsystem 110 is operating in a debug mode. Thus, at operation 316, the host system 120 may process the status information and perform one or more debug operations. For example, the debug operations may include requesting additional status information from the memory subsystem 110, sending debug commands or requests into various two-way ports that have been activated in the memory subsystem 110, requesting access to certain vendor-specific logs in the memory subsystem 110, providing a firmware update to fix an identified vulnerability, or other debug operations.
[0041] Figure 4 is a flow diagram of an example method of recovering an unresponsive memory subsystem using backup firmware enabled by a host-initiated hardware recovery signal according to some embodiments of the present disclosure. Method 400 may be performed by processing logic that may include hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executing on a processing device), or a combination thereof. In some embodiments, method 400 is performed by Figure 1 the recovery manager component 113. Although shown in a particular sequence or order, the order of the process may be modified unless otherwise specified. Thus, the illustrated embodiments should be understood only as examples, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. Additionally, one or more processes may be omitted in various embodiments. Thus, not every embodiment requires all of the processes. Other process flows are possible.
[0042] At operation 405, processing logic (e.g., the recovery manager component 113) executes a restart sequence in response to the occurrence of a power cycle event. In one embodiment, the host system 120 may initiate a power cycle event by sending a command to the power controller 225 to temporarily stop power supply to the memory subsystem 110 and then restore power to the memory subsystem 110 after a period of time. This power cycle event may be initiated in response to the memory subsystem 110 being unresponsive to the host system 120. For example, the power cycle event may include temporarily powering down the memory subsystem 110 by the host system 120 in response to the memory subsystem 110 being unresponsive for a period of time.
[0043] At operation 410, the processing logic samples a recovery signal received from the host system via the sideband interface 124 and determines whether the recovery signal is asserted. In one embodiment, the host system 120 asserts the recovery signal by driving a voltage (e.g., a voltage above a threshold) representing a logic "high" state on the sideband interface 124, which may be received at the memory subsystem 110 (e.g., at the GPIO pin to which the sideband interface 124 is connected). In one embodiment, the recovery signal is asserted by the sideband interface 124 that is separate from the primary interface (e.g., the PCIe bus 122) between the memory subsystem 110 and the host system 120.
[0044] In response to determining that the recovery signal is not asserted, at operation 415, the processing logic retrieves and loads the default firmware image of the memory subsystem and boots the memory subsystem in the default operation mode using the default firmware image. In one embodiment, the recovery manager 113 retrieves the default firmware image 212 from the database 210 and loads this default firmware. The default firmware enables a default operation mode that may be used when no major system failures or errors are detected. In the default operation mode, certain debugging capabilities are not available. For example, certain ports may not be accessible by the host system 120, certain event logs may be restricted, and so on.
[0045] In response to determining that the recovery signal is asserted, at operation 420, the processing logic retrieves and loads the backup firmware image of the memory subsystem. In one embodiment, the recovery manager 113 retrieves the backup firmware image 214 from the database 210 and loads this backup firmware. In one embodiment, the backup firmware is pre-installed on the memory subsystem 110 such that it can be used in the case of a major system failure or error that requires additional debugging capabilities.
[0046] At operation 425, the processing logic boots the memory subsystem in a debug operation mode using a backup firmware image. As described herein, the debug operation mode can be different from the default operation mode, such that certain debug functionality that is not normally available in the default operation mode becomes available. That is, the debug operation mode permits one or more debug functions that are not permitted in the default operation mode.
[0047] At operation 430, the processing logic provides status information to the host system 120. This status information can include an indication that the memory subsystem has been restarted and the backup firmware has been loaded, as well as other information. The status information can include an indication of a fault or error (if known) that has caused the memory subsystem 110 to become unresponsive. Receiving the status information can indicate to the host system 120 that the memory subsystem 110 is operating in the debug operation mode.
[0048] At operation 435, the processing logic performs one or more debug functions in response to a host request. In one embodiment, the one or more debug functions include unlocking one or more bidirectional communication ports in the memory subsystem 110 to allow debug operations to be performed. In one embodiment, the one or more debug functions include authorizing access to one or more vendor - specific logs in the memory subsystem 110. In one embodiment, the one or more debug functions include downloading updated firmware to resolve a fault in the memory subsystem 110.
[0049] Figure 5 An example machine of a computer system 500 is illustrated, within which a set of instructions can be executed to cause the machine to perform any one or more of the methodologies discussed herein. In some embodiments, the computer system 500 can correspond to a host system (e.g., Figure 1 the host system 120) that includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 1 the memory subsystem 110) or can be used to execute the operations of a controller (e.g., to execute an operating system to perform the operations corresponding to Figure 1 the recovery manager component 113). In alternative embodiments, the machine can be connected (e.g., networked) to other machines in a LAN, intranet, extranet, and / or the Internet. The machine can operate as a server or client machine in a client - server network environment, as a peer machine in a peer - to - peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.
[0050] The machine can be a personal computer (PC), tablet PC, set-top box (STB), personal digital assistant (PDA), cellular phone, network device, server, network router, switch, or bridge, or any machine capable of executing a set of instructions (sequentially or otherwise) that specify actions to be taken by the machine. Additionally, although a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0051] Example computer system 500 includes a processing device 502, a main memory 504 (such as read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), etc.), a static memory 506 (such as flash memory, static random access memory (SRAM), etc.), and a data storage system 518, which communicate with each other via a bus 530.
[0052] Processing device 502 represents one or more general-purpose processing devices, such as a microprocessor, central processing unit, or the like. More specifically, the processing device can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or multiple processors implementing a combination of instruction sets. Processing device 502 can also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. Processing device 502 is configured to execute instructions 526 for performing the operations and steps discussed herein. Computer system 500 can further include a network interface device 508 to communicate via a network 520.
[0053] Data storage system 518 can include a machine-readable storage medium 524 (also referred to as a computer-readable medium) on which is stored one or more sets of instructions 526 or software embodying any one or more of the methodologies or functions described herein. The instructions 526 can also reside, completely or at least partially, within main memory 504 and / or processing device 502 during execution by computer system 500, which also constitutes a machine-readable storage medium. Machine-readable storage medium 524, data storage system 518, and / or main memory 504 can correspond to Figure 1 memory subsystem 110.
[0054] In one embodiment, the instructions 526 include those for implementing corresponding to Figure 1instructions for the functionality of the recovery manager component 113. Although the machine-readable storage medium 524 is shown as a single medium in the example embodiment, the term "machine-readable storage medium" should be regarded as including a single medium or multiple media that store one or more sets of instructions. The term "machine-readable storage medium" should also be regarded as including any medium that is capable of storing or encoding a set of instructions for execution by a machine and that causes the machine to perform any one or more of the methodologies of the present disclosure. Thus, the term "machine-readable storage medium" should be regarded as including (but not limited to) solid-state memory, optical media, and magnetic media.
[0055] Some parts of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived as a self-consistent sequence of operations leading to a desired result. The operations are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0056] However, it should be borne in mind that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may relate to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the memories or registers or other such information storage systems of the computer system.
[0057] The present disclosure also relates to an apparatus for performing the operations herein. This apparatus may be specially constructed for the intended purpose, or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a computer-readable storage medium, such as (but not limited to) any type of disk (including floppy disks, optical disks, CD-ROMs, and magneto-optical disks), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of media suitable for storing electronic instructions, each coupled to a computer system bus.
[0058] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general-purpose systems may be used in conjunction with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized devices to perform the methods. The structure of various such systems will appear as set forth in the appended claims. Additionally, the present disclosure has been described without reference to any particular programming language. It should be understood that various programming languages may be used to implement the teachings of the present disclosure described herein.
[0059] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon that may be used to program a computer system (or other electronic device) to perform a process in accordance with the present disclosure. The machine-readable medium includes any mechanism for storing information in a form readable by a machine, such as a computer. In some embodiments, the machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium such as read only memory ("ROM"), random access memory ("RAM"), magnetic disk storage media, optical storage media, flash memory components, and the like.
[0060] In the foregoing description, embodiments of the present disclosure have been described with reference to specific example embodiments of the present disclosure. It should be understood that various modifications may be made to the present disclosure without departing from the broader spirit and scope of the embodiments of the present disclosure set forth in the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.
Claims
1. A memory subsystem comprising: Memory device; A processing device operatively coupled to the memory device to perform operations comprising: executing a restart sequence in response to occurrence of a power cycling event; determining whether a resume signal is asserted by the host system via the sideband interface; responsive to determining that the resume signal is asserted, retrieving and loading a backup firmware image of the memory subsystem; and The memory subsystem is booted in a debug mode of operation using the backup firmware image.
2. The memory subsystem of claim 1, wherein the power cycling event comprises a temporary power outage for a period of time initiated by the host system in response to the memory subsystem being unresponsive. 3 . The memory subsystem of claim 1 , wherein the resume signal is asserted via the sideband interface that is separate from a primary interface between the memory subsystem and the host system.
4. The memory subsystem of claim 1 , wherein the processing device is configured to perform operations further comprising: In response to determining that the resume signal is not asserted, retrieving and loading a default firmware image for the memory subsystem; and The memory subsystem is booted in a default operating mode using the default firmware image.
5. The memory subsystem of claim 4, wherein the debug operating mode permits one or more debug functions that are not permitted in the default operating mode.
6. The memory subsystem of claim 5, wherein the one or more debug functions include unlocking one or more bidirectional communication ports in the memory subsystem to allow debug operations to be performed.
7. The memory subsystem of claim 5, wherein the one or more debug functions include authorizing access to one or more vendor specific logs in the memory subsystem.
8. The memory subsystem of claim 5, wherein the one or more debug functions include downloading update firmware to resolve a fault in the memory subsystem.
9. A method comprising: executing a restart sequence of the memory subsystem in response to occurrence of a power cycling event; determining whether a resume signal is asserted by the host system via the sideband interface; responsive to determining that the resume signal is asserted, retrieving and loading a backup firmware image of the memory subsystem; and The memory subsystem is booted in a debug mode of operation using the backup firmware image.
10. The method of claim 9, wherein the power cycling event comprises a temporary power outage for a period of time initiated by the host system in response to the memory subsystem being unresponsive.
11. The method of claim 9, wherein the resume signal is asserted via the sideband interface that is separate from a primary interface between the memory subsystem and the host system.
12. The method according to claim 9, further comprising: in response to determining that the resume signal is not asserted, retrieving and loading a default firmware image for the memory subsystem; and The memory subsystem is booted in a default operating mode using the default firmware image.
13. The method of claim 12, wherein the debug mode of operation permits one or more debug functions that are not permitted in the default mode of operation.
14. The method of claim 13, wherein the one or more debug functions include unlocking one or more bidirectional communication ports in the memory subsystem to allow debug operations to be performed.
15. The method of claim 13, wherein the one or more debug functions include authorizing access to one or more vendor-specific logs in the memory subsystem.
16. The method of claim 13, wherein the one or more debug functions include downloading update firmware to resolve a fault in the memory subsystem.
17. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations comprising: executing a restart sequence of the memory subsystem in response to occurrence of a power cycling event; determining whether a resume signal is asserted by the host system via the sideband interface; responsive to determining that the resume signal is asserted, retrieving and loading a backup firmware image of the memory subsystem; and The memory subsystem is booted in a debug mode of operation using the backup firmware image.
18. The non-transitory computer-readable storage medium of claim 17, wherein the power cycling event comprises a temporary power outage for a period of time initiated by the host system in response to the memory subsystem being unresponsive.
19. The non-transitory computer-readable storage medium of claim 17, wherein the resume signal is asserted via the sideband interface that is separate from a primary interface between the memory subsystem and the host system.
20. The non-transitory computer-readable storage medium of claim 17, wherein the instructions cause the processing device to perform operations further comprising: In response to determining that the resume signal is not asserted, retrieving and loading a default firmware image for the memory subsystem; and The memory subsystem is booted in a default operating mode using the default firmware image, wherein the debug operating mode permits one or more debug functions not permitted in the default operating mode.