Memory Subsystem Die Fail Storm Mitigation via Abbreviated Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Memory subsystems face performance degradation and deadlock situations during a memory die fail storm due to failed memory dice, where error correction and data recovery processes consume resources and fail to keep up with demand, leading to a lack of free block stripes and unfulfilled host write requests.
Innovation Solution
Implementing an abbreviated error recovery procedure that skips read retry operations and relies solely on RAIN recovery when error rates exceed a threshold, ensuring efficient data transfer from failed to functioning dice, thereby reducing resource wastage and increasing the rate of data transfer, and maintaining free block stripes to prevent deadlocks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If full error recovery procedure with read retry operations is performed, then data accuracy is improved, but system performance and resource efficiency deteriorate during die fail storm
Solution Approach 1:
The patent applies partial action by implementing an abbreviated error recovery procedure that performs only RAIN recovery without read retry operations when error rates exceed a threshold. This selective approach performs only the necessary recovery action (RAIN) while omitting redundant actions (read retries), thereby improving system performance during die fail storms while still ensuring data accuracy through the RAIN mechanism.
Solution Approach 2:
The patent changes the error recovery parameter from full recovery (with read retries) to abbreviated recovery (without read retries) based on the error rate threshold. When the error rate exceeds the threshold indicating a die fail storm, the system switches to the abbreviated procedure, dynamically adjusting the recovery strategy to match the failure condition and optimize performance.
2Reliability
If read retry operations are performed on failed dice, then data recovery attempts are improved, but resource consumption and time waste increase
Solution Approach 1:
The patent extracts and removes the read retry operations from the error recovery procedure when operating in die fail storm conditions. By taking out the redundant read retry step and retaining only the essential RAIN recovery operation, the system eliminates time waste associated with repeated failed read attempts while preserving effective data recovery through RAIN.
Solution Approach 2:
The patent applies skipping by bypassing the read retry operations entirely when error rates indicate a die fail storm. The system rushes through the error recovery process by directly performing RAIN recovery without the intermediate read retry step, thereby reducing the time lost to unsuccessful recovery attempts while still achieving data recovery through the more efficient RAIN mechanism.
3Reliability
If error correction processes are intensified to handle failed dice, then data integrity is improved, but available free block stripes decrease leading to deadlocks
Solution Approach 1:
The patent applies partial action by performing only the essential RAIN recovery operation without the excessive read retry attempts. This selective approach maintains data integrity through RAIN while consuming fewer resources, thereby preserving free block stripe availability and preventing deadlocks that would occur with intensified full error correction processes.
4Reliability
If full error recovery procedure is used, then comprehensive data recovery is achieved, but resource efficiency and transfer rate decrease
Solution Approach 1:
The patent extracts the redundant read retry operations from the full error recovery procedure, retaining only the essential RAIN recovery step. This extraction maintains comprehensive data recovery capability through RAIN while eliminating the resource overhead and time consumption of read retries, thereby improving the data transfer rate during die fail storms.
Data Source
AI summary
A method is described that includes processing, by a memory subsystem, a read memory command that is addressed to a first die of a memory device. The memory subsystem determines whether processing the read memory command failed to correctly read user data from the first die and, in response to determining that processing the read memory command failed to correctly read user data from the first die, determines whether the first die has failed. In response to determining that the first die has failed, the memory subsystem performs an abbreviated error recovery procedure to successfully perform the read memory command instead of a full error recovery procedure.


