Boot Partition Crash Dump Storage in NVMe SSDs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In data storage systems, particularly SSDs with NVMe protocols, system crashes often occur due to firmware errors or data structure corruption, leading to challenges in storing and retrieving crash-dump information, as existing solutions require valid firmware to be installed before accessing crash-dump data, which can fail if the initial firmware is faulty or corrupted.
Innovation Solution
Storing crash-dump information in a boot partition area of the NVM device and/or a virtual boot partition area of the RAM, allowing retrieval by a host device without installing valid replacement firmware, thus bypassing the need for initial firmware integrity and ensuring data availability even if the NAND is faulty.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If crash-dump information is stored in traditional storage locations requiring firmware validation, then data integrity can be ensured, but retrieval fails when initial firmware is faulty or corrupted
Solution Approach 1:
The patent extracts crash-dump information storage from the traditional firmware-dependent storage locations and places it in the boot partition, which can be accessed independently of firmware status. This allows the host to retrieve crash-dump data directly from the boot partition without needing valid firmware installed, resolving the contradiction between data availability and retrieval simplicity.
Solution Approach 2:
The boot partition acts as an intermediary storage location between the crash-dump information and the host system. It serves as a mediator that enables data retrieval without requiring the intermediary firmware layer to be functional, thus allowing direct host access to crash-dump data while maintaining data integrity.
2Reliability
If valid replacement firmware is installed before accessing crash-dump data, then system stability is improved, but time is lost due to firmware installation requirement
Solution Approach 1:
The system performs preliminary action by storing crash-dump information in the boot partition during the crash event itself, before any firmware replacement is needed. This preliminary storage in an independently accessible location eliminates the need for time-consuming firmware installation before data retrieval, as the host can directly access the boot partition immediately.
3Speed
If crash-dump information is stored only in volatile RAM, then retrieval speed is improved, but data loss occurs upon power cycle
Solution Approach 1:
The patent merges the advantages of both volatile and non-volatile storage by implementing a dual-location storage system. Crash-dump information is stored in both the virtual boot partition in RAM (for fast retrieval) and the boot partition in NVM (for persistent storage). This combination ensures both rapid access and data retention across power cycles.
Solution Approach 2:
The system changes the storage parameter from exclusively volatile to a hybrid volatile-non-volatile configuration. By storing crash-dump data in both RAM and NVM boot partitions, the system transforms the storage characteristics to simultaneously achieve fast retrieval (from RAM) and persistent retention (from NVM).
Data Source
AI summary
The present disclosure describes technologies and techniques for use by a data storage controller—such as a controller for use with a NAND or other non-volatile memory (NVM)—to store crash-dump information in a boot partition following a system crash within the data storage controller. Within illustrative examples described herein, the boot partition may be read by a host device without the host first re-installing valid firmware into the data storage controller following the system crash. In the illustrative examples, the data storage controller is configured for use with versions of Peripheral Component Interconnect (PCI) Express—Non-Volatile Memory express (NVMe) that provide support for boot partitions in the NVM. The illustrative examples additionally describe virtual boot partitions in random access memory (RAM) for storing crash-dump information if the NAND has been corrupted, where the crash-dump information is retrieved from the RAM without power-cycling the RAM.


