Heterogeneous Memory Channel Recovery with Segmented ECC

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing memory systems face challenges in error recovery, particularly introducing fetch gaps and incurring additional latency, while also risking system hangs during long recovery sequences, which are not suitable for applications requiring low latency and high availability.

Innovation Solution

A memory system with a tiered error correction and recovery mechanism that allows for gapless fetches, separate channel marking for data integrity, and programmable timers to manage recovery processes, enabling one channel to recover without disrupting other channels, thus maintaining system performance and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If error recovery operations are performed on a failing memory channel, then the channel can be brought back into operational mode, but system operations may be hung for extended periods (e.g., 10 ms for clock recalibration and PLL locking)

Engineering Contradiction:
Improvememory channel availabilityVSAvoidsystem hang duration
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The memory system divides the memory channel into multiple quarters, allowing independent error detection and recovery for each quarter. This segmentation enables partial channel recovery without requiring complete channel shutdown, reducing system hang duration while maintaining reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary error detection using CRC codes before full recovery operations are initiated. This allows early identification of failing quarters and enables targeted recovery actions, preventing unnecessary system-wide hangs and reducing recovery time.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If the entire data line is delayed until all ECC is clean during error recovery, then fetch gaps are avoided, but undue latency is introduced on all lines including those without errors

Engineering Contradiction:
Improvedata integrityVSAvoidfetch latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The data line is divided into multiple quarters with independent error status tracking. Only quarters with detected errors are delayed for recovery, while error-free quarters are forwarded immediately. This selective handling eliminates unnecessary latency on clean data segments while maintaining data integrity through targeted error recovery.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different recovery actions are applied to different quarters of the data line based on their individual error status. Quarters with errors undergo recovery procedures while quarters without errors are processed normally, creating localized quality control that optimizes both reliability and latency performance.

Inventive Principle:
Principle #3Local quality

3Reliability

If separate address/protocol tags are maintained for each quarter line during error recovery, then fetch gaps can be avoided, but device complexity increases

Engineering Contradiction:
Improvecontinuous data flowVSAvoidaddress tagging mechanism
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system merges the address tagging mechanism with the existing quarter-line error status information. Instead of maintaining completely separate address tags for each quarter, the error status bits are integrated into the existing address protocol, reducing the additional complexity while still enabling gapless fetches through selective quarter-line handling.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS8775858B2Heterogeneous recovery in a redundant memory system
Publication Date: 2014.07.08 GLOBALFOUNDRIES US INC
  • US8775858B2 patent drawing
  • US8775858B2 patent drawing
  • US8775858B2 patent drawing

AI summary

Providing heterogeneous recovery in a redundant memory system that includes a memory controller, a plurality of memory channels in communication with the memory controller, an error detection code mechanism configured for detecting a failing memory channel, and an error recovery mechanism. The error recovery mechanism is configured for receiving notification of the failing memory channel, for performing a recovery operation on the failing memory channel while other memory channels are performing normal system operations, for bringing the recovered channel back into operational mode with the other memory channels for store operations, for continuing to mark the recovered channel to guard against stale data, for removing any stale data after the recovery operation is complete, and for removing the mark on the recovered channel to allow the normal system operations with all of the memory channels, the removing based on the removing any stale data being complete.