Fault-Tolerant Memory Blade With Dual-Path CXL Replication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing reliance on external memory disaggregation in data centers creates a significant risk of single points of failure, leading to substantial disruptions due to the blast radius of memory module failures, necessitating fault-tolerant memory solutions with higher reliability and availability.
Innovation Solution
Implementing a fault-tolerant memory system using CXL switches for dual-path data replication and error correction, enabling clients to write data to dual-ported memory modules through redundant fabric pathways, ensuring continuous access and bandwidth even in the event of failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If external memory disaggregation is implemented to increase memory capacity and flexibility, then memory scalability is improved, but system reliability deteriorates due to single points of failure
Solution Approach 1:
The system segments memory access paths into multiple independent fabric switches and pathways. Each memory module can be accessed through different fabric switches, creating segmented access paths that isolate failures to specific segments rather than affecting the entire memory system.
Solution Approach 2:
The system implements redundant fabric pathways and dual-ported memory modules before failures occur. This beforehand cushioning ensures that alternative paths are already in place and ready to handle traffic immediately when a failure occurs, preventing service disruption.
2Reliability
If redundant fabric pathways are implemented to improve fault tolerance, then system reliability is improved, but device complexity increases
Solution Approach 1:
The fabric switches are designed with universal connectivity capabilities, where each switch can serve multiple memory modules and multiple clients simultaneously. The dual-ported memory modules also provide multi-functionality by accepting writes from either fabric pathway, reducing the need for dedicated redundant components.
Solution Approach 2:
The system creates copies of data in dual-ported memory modules accessible through different fabric pathways. Instead of creating complex redundant hardware paths, the system uses data copying to memory locations that are already accessible through multiple paths, simplifying the overall architecture.
3Reliability
If dual-casting data to paired memory modules is implemented to ensure redundancy, then data availability is improved, but bandwidth requirements increase
Solution Approach 1:
The system performs preliminary data placement in dual-ported memory modules during normal operation, so that redundant copies are already in place before failures occur. This eliminates the need for data replication after failures, reducing the bandwidth overhead associated with emergency data copying.
Solution Approach 2:
The system merges the redundancy function into the existing memory architecture by using dual-ported memory modules that can be accessed through either fabric pathway. This combines the redundant storage function with the memory access function, avoiding the need for separate redundant data paths that would consume additional bandwidth.
Data Source
AI summary
The technology disclosed herein provides a memory blade including a plurality of fabric switches configured to receive commands from a plurality of host clients, an address decoder and tracker circuit communicatively connected to the fabric switches and configured to determine the source of commands received at the fabric switches an aggregator crossbar configured to provide bandwidth aggregation between host clients and a plurality of memory modules; and a buffer module configured to couple the commands from the plurality of host clients with the plurality of memory modules.


