Disaggregated Memory Access via InfiniBand Transaction Datapaths
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Datacenters face inefficient memory usage due to non-disaggregated main system memory, leading to stranded resources and inefficient utilization of processing resources, especially when CPUs are fully utilized.
Innovation Solution
A system and method for memory disaggregation using a software-defined hardware datapath that enables remote mastering of system interconnect transactions over InfiniBand, leveraging DPU/IB NIC products to facilitate dynamic configuration and centralized control of link establishment between remote resources, allowing flexible and efficient use of memory resources across the datacenter network.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If local memory resources are used by processing resources, then memory access speed is improved, but memory utilization efficiency deteriorates due to stranded resources
Solution Approach 1:
The patent segments memory resources from processing resources, creating separate memory pools that can be independently managed and allocated. Memory is divided into discrete allocable units that can be dynamically assigned to different processing resources based on demand, rather than being permanently bound to specific CPUs.
Solution Approach 2:
The patent introduces a fabric interconnect (InfiniBand network) as an intermediary between processing resources and memory resources. This intermediary enables remote memory access while maintaining efficient communication pathways, allowing CPUs to access memory pools located in different physical locations without direct local coupling.
2Productivity
If memory is disaggregated across the network, then memory utilization efficiency is improved, but access latency increases
Solution Approach 1:
The fabric interconnect serves as a specialized intermediary optimized for low-latency communication. Unlike general-purpose networks, the InfiniBand fabric provides direct routed paths with predictable performance characteristics, minimizing the time penalty associated with remote memory access.
Solution Approach 2:
The system dynamically selects memory locations and access paths based on current workload requirements and network conditions. Memory pools can be dynamically allocated and deallocated, and the fabric routing can adapt to optimize data flow, reducing overall access latency through intelligent resource management.
3Adaptability or versatility
If remote memory access is enabled, then adaptability of memory resources is improved, but system complexity increases
Solution Approach 1:
The fabric interconnect provides universal communication capabilities that work with multiple types of processing resources and memory configurations. The same infrastructure supports various memory pool configurations, different CPU architectures, and multiple workload types, reducing the need for specialized components for each scenario.
Solution Approach 2:
The system implements feedback mechanisms where the fabric and memory management infrastructure continuously monitor resource usage, access patterns, and system state. This feedback enables automated resource allocation, load balancing, and optimization decisions that simplify management while maintaining high adaptability to changing requirements.
Data Source
AI summary
A system comprises a first processing block configured to receive, from a first local resource, a formatted transaction in a format that is not recognizable by a remote endpoint; determine a first transaction category, from among a plurality of transaction categories, of the formatted transaction based on content of the formatted transaction; perform one or operations on the formatted transaction based on the first transaction category to form a reformatted transaction in a format that is recognizable by the remote endpoint; and place the reformatted transaction in a queue for transmission to the remote endpoint.


