Kernel Masking DRAM Defects via ECC and Bad Page Retirement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
As DRAM process technology scales down, there is a challenge in maintaining data reliability due to decreased cell capacitance and increased cell transistor leakage, leading to variance in cell retention and the need for frequent refresh or error correction, which impacts memory bandwidth and power consumption.
Innovation Solution
A system and method for kernel masking DRAM defects, involving an ECC module and a bad page masking module that detect and correct errors, store error data in a non-volatile memory, and retire kernel pages if error counts exceed a threshold, thereby masking defects and maintaining data integrity without significant increases in refresh frequency or silicon area usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If DRAM process technology is scaled down to increase memory capacity, then memory density is improved, but cell retention reliability deteriorates due to decreased cell capacitance and increased transistor leakage
Solution Approach 1:
The patent segments the DRAM memory space into multiple banks, with each bank independently managed and monitored for errors. This segmentation allows the system to isolate and handle defects in specific regions without affecting the entire memory array, thereby maintaining reliability while scaling down cell size to increase overall capacity.
Solution Approach 2:
The patent introduces an intermediary error detection and correction mechanism that monitors cell retention and intervenes when errors occur. This intermediary layer between the physical cells and the logical memory space allows the system to maintain data integrity even as cell retention characteristics deteriorate due to scaling.
2Reliability
If refresh frequency is increased to maintain data reliability in scaled-down DRAM cells, then cell retention reliability is improved, but memory bandwidth deteriorates and standby power consumption increases
Solution Approach 1:
The patent applies local quality by implementing error correction and monitoring only where needed - specifically in scaled-down cells that exhibit retention problems. Rather than uniformly increasing refresh frequency across the entire memory array, the system selectively applies enhanced error handling to problematic regions, thereby maintaining reliability without proportionally increasing bandwidth consumption and power usage.
3Reliability
If block error correction is implemented to correct multiple simultaneous errors, then data reliability is improved, but silicon area increases significantly
Solution Approach 1:
The patent applies partial action by implementing error correction capabilities that are sufficient for the actual error rates observed in scaled-down DRAM, rather than designing for worst-case scenarios. The system uses parity bits and error detection codes that can handle the typical single-bit errors that occur, without over-engineering for rare multiple simultaneous errors, thereby achieving adequate reliability with minimal silicon area overhead.
Data Source
AI summary
Systems, methods, and computer programs are disclosed for kernel masking dynamic random access memory (DRAM) defects. One such method comprises: detecting and correcting a single-bit error associated with a physical address in a dynamic random access memory (DRAM); receiving error data associated with the physical address from the DRAM; storing the received error data in a failed address table located in a non-volatile memory; and retiring a kernel page corresponding to the physical address if a number of errors associated with the physical address exceeds an error count threshold.


