Disaggregated Memory Pool With Tiered Latency Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computing systems face challenges in managing varying memory requirements and latency needs across applications and hosts, particularly in shared memory pools where different applications and hosts have different memory quantity and performance demands.
Innovation Solution
A scalable, disaggregated memory pool architecture utilizing a multi-dimensional array of memory nodes connected via a hyper torus switching fabric, allowing for latency management and access to memory nodes based on application-specific requirements, with each node potentially having computational processing circuits and connected through CXL connections.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If memory nodes are connected in a shared memory pool to increase memory capacity, then the quantity of memory is improved, but the latency varies and cannot be optimized for different applications
Solution Approach 1:
The memory pool is segmented into multiple memory nodes organized in a hierarchical structure with root memory nodes and leaf memory nodes. Each node can be independently accessed, allowing applications to select specific nodes based on latency requirements. This segmentation enables different portions of the memory pool to serve different performance needs simultaneously.
Solution Approach 2:
The patent introduces a hierarchical dimension to the memory pool architecture, organizing memory nodes in multiple levels (root nodes at level 0, leaf nodes at level N). This dimensional organization allows the system to provide differentiated access paths, where some applications can access root nodes for low-latency operations while others access leaf nodes for higher capacity operations.
2Productivity
If a unified memory pool is used to serve multiple hosts, then memory utilization is improved, but different applications cannot receive customized latency performance
Solution Approach 1:
Different regions of the memory pool (different memory nodes) are assigned different performance characteristics. Root memory nodes provide low-latency access for performance-critical applications, while leaf memory nodes provide high-capacity access for less time-sensitive workloads. This local quality differentiation allows the unified pool to serve diverse performance needs.
Solution Approach 2:
The system dynamically assigns memory nodes to applications based on their performance requirements. The hierarchical structure enables flexible allocation where applications can be directed to access specific nodes or ranges of nodes appropriate to their latency needs, allowing the memory pool to adapt to varying workload requirements in real-time.
3Quantity of substance
If memory nodes are added to increase capacity, then the quantity of memory is improved, but the system complexity increases
Solution Approach 1:
The hierarchical memory pool structure serves multiple functions simultaneously: it provides capacity expansion through additional nodes, enables latency differentiation through selective access paths, and maintains manageable complexity through standardized node interfaces. The same hierarchical framework supports both performance optimization and capacity scaling.
Solution Approach 2:
The memory pool is organized as nested hierarchical levels where root nodes contain or manage access to subordinate leaf nodes. This nesting allows the system to scale by adding complete node subsets at appropriate levels without fundamentally restructuring the entire system, thereby managing complexity while increasing capacity.
Data Source
AI summary
A system including a scalable memory pool. In some embodiments, the system includes: a first memory node, including a first memory; a second memory node, including a second memory; and a memory node switching fabric connected to the first memory node and the second memory node, the memory node switching fabric being configured to provide access, via the first memory node: with a first latency, to the first memory, and with a second latency, greater than the first latency, to the second memory.


