Fault Tolerant Memory System with Dynamic RAS Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current memory systems in data processing apparatuses face limitations in reliability, availability, and serviceability due to the inability to effectively handle bit errors, memory node controller failures, and multi-bit errors, while also being inefficient in terms of resource utilization and dynamic adaptation to changing component performance.

Innovation Solution

The implementation of a memory system with an RAS Management Unit that employs redundancy, error correction codes, and dynamic data placement and allocation strategies, including the use of criticality bits to identify and protect critical data, and secondary storage and memory node controllers to ensure fault tolerance and adapt to changing conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If redundant storage is used to tolerate bit errors, then reliability is improved, but device complexity and cost increase due to duplicating entire memory systems

Engineering Contradiction:
Improvefault toleranceVSAvoidsystem duplication
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the memory system into multiple memory nodes, each independently manageable. Instead of duplicating the entire system, individual nodes can be isolated and replaced. This segmentation allows partial redundancy without full system duplication, reducing complexity while maintaining fault tolerance.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses copying by maintaining duplicate copies of data within the same memory system through error correction codes and redundancy mechanisms. Rather than duplicating entire memory systems, data copies are created selectively, reducing device complexity while preserving reliability.

Inventive Principle:
Principle #26Copying

2Reliability

If error correction codes are used to correct bit errors, then reliability is improved, but the capability is limited to a fixed number of correctable bits

Engineering Contradiction:
Improveerror correctionVSAvoiderror correction capacity
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic error correction by allowing the system to adapt its correction capacity based on actual error conditions. The memory management unit can switch between different error correction strategies and allocate resources dynamically, enabling the system to handle varying numbers of bit errors beyond fixed ECC limitations.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes parameters by allowing flexible configuration of error correction capabilities. The system can adjust the number of parity bits, switch between different ECC schemes, and modify correction thresholds based on observed error patterns, thereby increasing adaptability to different error scenarios.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If static system configuration is used, then device complexity is reduced, but the system cannot adapt to changing performance of memory and storage components

Engineering Contradiction:
Improvesystem configurationVSAvoiddynamic adaptation
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent makes the system configuration dynamic by implementing a memory management unit that continuously monitors component health and performance. The system can dynamically reconfigure data placement, migrate data between memory nodes, and adjust resource allocation based on real-time conditions, enabling adaptation without excessive complexity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent incorporates feedback mechanisms where the memory management unit monitors the health and performance of memory and storage components. This feedback enables the system to detect degradation, predict failures, and proactively reconfigure resources, maintaining optimal performance while adapting to changing component conditions.

Inventive Principle:
Principle #23Feedback

4Device complexity

If existing data placement methods are used, then device complexity is reduced, but fault tolerance is insufficient for memory node controller failures and multi-bit errors

Engineering Contradiction:
Improvedata placementVSAvoidfault tolerance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent segments data across multiple memory nodes and controllers, distributing copies strategically. This segmentation enables the system to tolerate failures of individual nodes or controllers while maintaining data availability. The segmented approach provides enhanced fault tolerance without requiring complex centralized management.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10884850B2Fault tolerant memory system
Publication Date: 2021.01.05 ARM LTD
  • US10884850B2 patent drawing
  • US10884850B2 patent drawing
  • US10884850B2 patent drawing

AI summary

A memory system for a data processing apparatus includes a fault management unit, a memory controller (such as a memory management unit or memory node controller), and one or more storage devices accessible via the memory controller and configured for storing critical data. The fault management unit detects and corrects a fault in the stored critical data, a storage device or the memory controller. A data fault may be corrected using a copy of the data, or an error correction code, for example. A level of failure protection for the critical data, such as a number of copies, an error correction code or a storage location in the one or more storage devices, is determined dependent upon a failure characteristic of the device. A failure characteristic, such as an error rate, may be monitored and updated dynamically.