Zone Memory Read-Failure Recovery for Cache and Non-Cache Blocks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional approaches for handling block read failure in zone-based memory sub-systems are insufficient, leading to data loss or corruption during operations like cache-to-non-cache data migration, cache block refresh, or when a controller finishes a partially written zone, especially in NAND-type memory devices.
Innovation Solution
The memory sub-system employs enhanced mechanisms to handle block read failures by relocating data on detection of SLC or QLC Uncorrectable Error Code Correction (UECC) and implementing quick recovery processes, maintaining data integrity and reducing downtime, particularly in solid-state drives (SSDs) using a zone architecture like 3S/1Q.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional block read failure handling is used in zone-based memory sub-systems, then device complexity is reduced, but data integrity deteriorates due to data loss or corruption during cache-to-non-cache migration, cache block refresh, or zone completion operations
Solution Approach 1:
The system performs preliminary actions by completing ongoing cache block programming operations and queued programs before handling read failures. This ensures that data is fully written to non-cache blocks before attempting to read from cache blocks, preventing data loss during migration and refresh operations
Solution Approach 2:
The controller acts as an intermediary between cache blocks and non-cache blocks, managing the read failure handling process by coordinating data relocation operations. The controller monitors read failures, determines appropriate recovery actions, and executes them while maintaining zone consistency
2Reliability
If enhanced block read failure handling mechanisms are implemented, then data integrity is improved through data relocation on UECC detection, but device complexity increases due to additional monitoring and recovery processes
Solution Approach 1:
The system continuously monitors read operations for uncorrectable errors (UECC) and uses this feedback to trigger appropriate recovery actions. When a read failure is detected, the system feedbacks by initiating data relocation from cache to non-cache blocks, ensuring data integrity is maintained
Solution Approach 2:
The memory sub-system performs self-service by automatically detecting read failures and executing recovery operations without external intervention. The controller autonomously monitors for UECC, determines data relocation needs, and executes recovery processes independently
3Reliability
If quick recovery processes are implemented for block read failures, then system reliability is improved with minimal downtime, but productivity may be affected due to additional operations required during failure recovery
Solution Approach 1:
The system performs preliminary actions by completing ongoing cache block programming operations and queued programs before handling read failures. This ensures that data is fully written to non-cache blocks before attempting to read from cache blocks, preventing data loss during migration and refresh operations
Solution Approach 2:
The recovery process maintains continuity by executing data relocation operations during zones that are marked as finished, allowing recovery operations to proceed without interrupting normal data access paths. The system continues useful actions while performing necessary recovery operations in parallel
Data Source
AI summary
Various embodiments provide handling block read failure in a memory sub-system that supports zones. In particular, some embodiments described herein handle block read failure during a data read (e.g., host data write) of a cache block or a non-cache block of a zone on a memory device on a memory sub-system, block read failure during refresh of a cache block or a non-cache block of a zone on a memory device on a memory sub-system, block read failure during migration of data between a cache block and a non-cache block of a zone on a memory device on a memory sub-system, or some combination thereof.


