Memory Controller Parity-Group Assignment for Multi-Mode Data Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data storage systems face challenges in efficiently recovering data due to multiple failure modes in memory devices, such as WL-to-WL shorts and WL-to-substrate leakage, which are not effectively addressed by conventional RAID schemes, leading to high storage space and cost requirements.
Innovation Solution
A controller is designed to assign memory cell-groups to parity-groups in a way that no two cell-groups belong to the same WL or adjacent WLs in the same section, allowing for data recovery using remaining cell-groups in case of specific failure modes, reducing the number of parity-groups and redundancy storage space needed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional RAID schemes are used for data storage, then data recovery capability is provided, but storage space and cost requirements increase significantly
Solution Approach 1:
The memory device is segmented into multiple sections, with each section containing multiple word lines. The patent assigns cell-groups from different sections to the same parity-group, creating a hierarchical structure that enables targeted recovery. When a failure occurs in one section, only the affected cell-groups need recovery rather than entire parity-groups, reducing the storage overhead required for redundancy.
Solution Approach 2:
The patent implements failure-mode-specific recovery strategies by analyzing the pattern of failed word lines. Different recovery procedures are applied depending on whether the failure is localized to a single section or affects multiple sections. This localized approach allows the system to use minimal redundancy for common failure modes while maintaining robustness for rarer patterns, optimizing the balance between reliability and storage efficiency.
2Reliability
If redundancy storage is increased to handle multiple failure modes, then data recovery reliability improves, but storage efficiency decreases
Solution Approach 1:
The system dynamically selects recovery procedures based on the detected failure mode. The controller analyzes the pattern of failed word lines and adaptively chooses the appropriate recovery strategy from multiple options. This dynamic approach allows the system to maintain high reliability for all failure modes while optimizing storage efficiency by using the minimal necessary redundancy for each specific failure scenario rather than over-provisioning for all possible failures simultaneously.
Solution Approach 2:
The patent changes the parameters of the recovery process based on the failure mode detected. Different recovery procedures with varying levels of redundancy and computational complexity are applied depending on the specific failure pattern. This parameter adaptation allows the system to achieve high reliability for multiple failure modes without consistently using excessive storage space, as the redundancy level is adjusted to match the actual failure scenario.
3Reliability
If more parity-groups are created to cover all failure modes, then comprehensive data protection is achieved, but device complexity increases
Solution Approach 1:
The patent creates a universal recovery framework where a single parity-group structure serves multiple failure modes. By organizing cell-groups from different sections into parity-groups with specific assignment rules, the same parity-group can recover different types of failures depending on which cell-groups are affected. This multi-functional approach eliminates the need for separate parity-groups for each failure mode, reducing overall system complexity while maintaining comprehensive protection.
Solution Approach 2:
The system implements a feedback mechanism where the controller detects the actual failure mode and uses this information to select the appropriate recovery procedure. The failure detection feedback allows the system to activate only the necessary recovery paths rather than maintaining complex redundant structures for all possible failures. This feedback-driven approach simplifies the device architecture by enabling on-demand recovery rather than pre-configuring for every potential failure scenario.
Data Source
AI summary
A controller includes an interface and a processor. The interface is configured to communicate with a memory including multiple memory cells organized in at least two sections each including multiple sets of word lines (WLs), wherein in a first failure mode multiple WLs fail in a single section, and in a second failure mode a WL fails in multiple sections. The processor is configured to assign multiple cell-groups of the memory cells to a parity-group, such that (i) no two cell-groups in the parity-group belong to a same WL, and (ii) no two cell-groups in the parity-group belong to adjacent WLs in a same section, and, upon detecting a failure to access a cell-group in the parity-group, due to either the first or second failure modes but not both failure modes occurring simultaneously, to recover the data stored in the cell-group using one or more remaining cell-groups in the parity-group.


