SSD Failure Information Reporting via Log Page 0x31
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems fail to provide useful failure information from solid state drives (SSDs) to host devices, as the existing log information does not contain relevant failure data, leading to ineffective error handling and recovery in SSDs during Native Command Queuing (NCQ) operations.
Innovation Solution
A system and method for sending failure information from SSDs to host devices, which includes detecting errors during SSD operations, receiving a command for failure information, and transmitting relevant failure data, such as error location, data integrity, and persistent or transient failure status, using log pages or other formats, enabling effective error recovery and management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If existing log information is used for error reporting in SATA systems, then the drive can return basic command failure information, but the information is not useful for SSD-specific failure analysis and recovery
Solution Approach 1:
The patent introduces new parameters and data structures specifically designed for SSD failure reporting. It defines new log page formats (such as Log Page 0x31) that include SSD-specific failure information parameters like media error status, wear level indicators, and block replacement information, transforming the generic log structure into an SSD-optimized format that provides meaningful failure data.
Solution Approach 2:
The patent segments the error information into distinct categories and fields within the log structure. It divides failure information into separate segments including command-specific error codes, media status indicators, and recovery status fields, allowing the host to process different types of failure information independently and efficiently.
2Reliability
If the drive stops all activity upon error occurrence in NCQ operations, then error propagation is prevented, but productivity and system responsiveness deteriorate
Solution Approach 1:
The patent implements dynamic error handling where the drive's operational state is adjusted based on the severity and type of error detected. Instead of a static stop-all-activity response, the system dynamically determines whether to continue operations, throttle performance, or halt specific queues based on real-time error conditions, maintaining productivity while ensuring reliability.
Solution Approach 2:
The patent establishes a feedback mechanism where the drive continuously monitors error conditions and communicates status to the host through updated log pages. This feedback loop allows the host to make informed decisions about command queuing and error recovery strategies, enabling the system to adapt to error conditions without complete activity cessation.
Data Source
AI summary
A system, method, and computer program product are provided for a host device to request and obtain failure information from a solid state drive (SSD). In operation, an error is detected during an operation associated with a solid state drive. Additionally, a command to return failure information is provided to the solid state drive by a host device. Further, the failure information is sent from the solid state drive to the host device, the failure information including failure information associated with the solid state drive.


