Fault Tolerant Memory System with Dynamic RAS Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current memory systems in data processing apparatuses face limitations in reliability, availability, and serviceability due to the inability to effectively handle bit errors, memory node controller failures, and multi-bit errors, while also being inefficient in terms of resource utilization and dynamic adaptation to changing component performance.
Innovation Solution
The implementation of a memory system with an RAS Management Unit that employs redundancy, error correction codes, and dynamic data placement and allocation strategies, including the use of criticality bits to identify and protect critical data, and secondary storage and memory node controllers to ensure fault tolerance and adapt to changing conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If redundant storage is used to tolerate bit errors, then reliability is improved, but device complexity and cost increase due to duplicating entire memory systems
Solution Approach 1:
The patent segments the memory system into multiple memory nodes, each independently manageable. Instead of duplicating the entire system, individual nodes can be isolated and replaced. This segmentation allows partial redundancy without full system duplication, reducing complexity while maintaining fault tolerance.
Solution Approach 2:
The patent uses copying by maintaining duplicate copies of data within the same memory system through error correction codes and redundancy mechanisms. Rather than duplicating entire memory systems, data copies are created selectively, reducing device complexity while preserving reliability.
2Reliability
If error correction codes are used to correct bit errors, then reliability is improved, but the capability is limited to a fixed number of correctable bits
Solution Approach 1:
The patent implements dynamic error correction by allowing the system to adapt its correction capacity based on actual error conditions. The memory management unit can switch between different error correction strategies and allocate resources dynamically, enabling the system to handle varying numbers of bit errors beyond fixed ECC limitations.
Solution Approach 2:
The patent changes parameters by allowing flexible configuration of error correction capabilities. The system can adjust the number of parity bits, switch between different ECC schemes, and modify correction thresholds based on observed error patterns, thereby increasing adaptability to different error scenarios.
3Device complexity
If static system configuration is used, then device complexity is reduced, but the system cannot adapt to changing performance of memory and storage components
Solution Approach 1:
The patent makes the system configuration dynamic by implementing a memory management unit that continuously monitors component health and performance. The system can dynamically reconfigure data placement, migrate data between memory nodes, and adjust resource allocation based on real-time conditions, enabling adaptation without excessive complexity.
Solution Approach 2:
The patent incorporates feedback mechanisms where the memory management unit monitors the health and performance of memory and storage components. This feedback enables the system to detect degradation, predict failures, and proactively reconfigure resources, maintaining optimal performance while adapting to changing component conditions.
4Device complexity
If existing data placement methods are used, then device complexity is reduced, but fault tolerance is insufficient for memory node controller failures and multi-bit errors
Solution Approach 1:
The patent segments data across multiple memory nodes and controllers, distributing copies strategically. This segmentation enables the system to tolerate failures of individual nodes or controllers while maintaining data availability. The segmented approach provides enhanced fault tolerance without requiring complex centralized management.
Data Source
AI summary
A memory system for a data processing apparatus includes a fault management unit, a memory controller (such as a memory management unit or memory node controller), and one or more storage devices accessible via the memory controller and configured for storing critical data. The fault management unit detects and corrects a fault in the stored critical data, a storage device or the memory controller. A data fault may be corrected using a copy of the data, or an error correction code, for example. A level of failure protection for the critical data, such as a number of copies, an error correction code or a storage location in the one or more storage devices, is determined dependent upon a failure characteristic of the device. A failure characteristic, such as an error rate, may be monitored and updated dynamically.


