Server Memory ECC for DRAM Chip Failure and Double-Bit Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing memory systems face challenges in providing reliable error correction and detection with minimal cost, especially as RAM device geometries shrink and multi-bit errors become more prevalent, leading to increased susceptibility to data corruption.

Innovation Solution

A novel error correction and detection system using 16 ECC bits to correct single-bit and multi-bit errors within a 128-bit data word, employing a unique ECC table and RAM error definition table for syndrome decode, which detects and corrects errors across multiple RAM devices while reducing the number of required RAM devices and ECC bits.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple SBC/DBD fields are created across the data word with separate check bits for each bit, then error correction capability is improved, but the cost of additional RAM devices increases significantly

Engineering Contradiction:
Improveerror correction capabilityVSAvoidnumber of RAM devices
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple separate SBC/DBD fields into a unified error correction scheme where 16 check bits protect the entire 128-bit data word. Instead of having separate check bit groups for different portions of the data, the invention combines all error correction functionality into a single integrated system that achieves familial 1-4 bit correction and non-familial double bit detection across the whole word.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The 16 check bits serve multiple functions simultaneously: they provide single bit correction, familial multi-bit correction (1-4 bits within the same RAM device), and non-familial double bit detection. This multi-functionality eliminates the need for separate specialized check bit groups for different error types, reducing the total number of required RAM devices.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Device complexity

If 2 ECC fields with 2-bit adjacency correction are used, then the cost in RAM devices is reduced, but not all two-bit errors across multiple devices are detected

Engineering Contradiction:
Improvenumber of RAM devicesVSAvoiderror detection capability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent changes the parameters of the ECC system by using 16 check bits organized to provide familial correction for 1-4 bit errors and non-familial detection for 2 bit errors. This parameter configuration allows the system to detect all two-bit errors across multiple devices while maintaining cost efficiency, overcoming the limitation of traditional 2-bit adjacency correction schemes.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If RAM device geometries are shrunk to increase density, then memory capacity is improved, but multi-bit error susceptibility increases

Engineering Contradiction:
Improvememory densityVSAvoidmulti-bit error susceptibility
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent implements beforehand cushioning by using 16 check bits that are calculated and stored in advance with the data word. These check bits are prepared beforehand to detect and correct errors that may occur during storage, providing a safety buffer against multi-bit errors that become more likely due to shrunk geometries. The familial correction capability specifically addresses multi-bit errors within the same device.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS7634709B2Familial correction with non-familial double bit error detection
Publication Date: 2009.12.15 UNISYS CORP
  • US7634709B2 patent drawing
  • US7634709B2 patent drawing
  • US7634709B2 patent drawing

AI summary

Error correction and error detection related to DRAM chip failures, particularly adapted server memory subsystems. It uses ×4 bit DRAM devices organized in a code word of 128 data bit words and 16 check bits. These 16 check bits are generated in such a way as to provide a code capable of 4 bit adjacent error correction within a family (i.e., in a ×4 DRAM) and double bit non-adjacent error detection across the entire 128 bit word, with single bit correction across the word as well. Each device can be thought of as a separate family of bits, errors occurring in more than one family are not correctable, but may be detected if only one bit in each of two families is in error. Syndrome generation and regeneration are used together with a specific large code word. Decoding the syndrome and checking it against the regenerated syndrome yield data sufficient for providing the features described.