Pluggable Compute Modules with NVMe Storage via Non-Transparent Bridge
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current networked computing environments face challenges in effectively utilizing the high bandwidth and low latency of NVMe storage appliances due to network throughput limitations, leading to bottlenecks and underutilization of resources, especially in applications requiring shared access to data across multiple computing nodes.
Innovation Solution
The implementation of pluggable compute and accelerator modules with non-transparent bridge (NTB) technology, which enables dynamic creation of virtual tunnels and proxy endpoints for secure, on-demand access to resources like NVMe storage, FPGAs, GPUs, and network adapters, using Field Programmable Gate Arrays (FPGAs) or ASICs, allowing peer-to-peer communication and efficient resource partitioning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If NVMe storage devices are connected through network fabrics (Ethernet, Fibre Channel), then shared access to storage across multiple computing nodes is enabled, but latency increases and throughput decreases compared to direct PCIe connection
Solution Approach 1:
The patent introduces a network fabric as an intermediary medium that enables shared access to NVMe storage devices across multiple computing nodes. The fabric acts as a mediator that transports NVMe commands and data between hosts and storage devices, allowing versatile shared access while managing the inherent performance trade-offs of network-based communication.
2Speed
If storage is directly attached to each computing node, then high bandwidth and low latency are achieved, but cost increases and shared access becomes less efficient
Solution Approach 1:
The patent makes NVMe storage devices universally accessible to multiple computing nodes through the network fabric, allowing a single storage device to serve multiple hosts simultaneously. This multi-functionality enables efficient shared access while maintaining the high-performance NVMe interface, eliminating the need for separate direct-attached storage for each node.
3Productivity
If network throughput is increased to accommodate NVMe performance, then more data can be transferred, but the fundamental network latency issue persists and cost increases
Solution Approach 1:
The patent changes the operational parameters of the network fabric by implementing specific optimizations including buffer management, interrupt coalescing, and command queuing mechanisms. These parameter changes allow the network fabric to handle NVMe workloads more efficiently, improving throughput while minimizing the impact of latency through intelligent data buffering and batch processing.
Data Source
AI summary
Methods and systems for a networked computing system are provided. One method includes generating, based on a first topology, a first proxy endpoint by a first device of a first pluggable compute module; establishing a communication tunnel between the first proxy endpoint and a non-volatile memory express (NVMe) storage device for peer-to-peer communication between the first proxy endpoint of the first device and a controller of the NVMe storage device. An NVMe translation module receives a request for the NVMe storage device from the first proxy endpoint and the NVMe translation module translates the request to an NVMe request for the NVMe storage device for accessing storage space at the NVMe storage device. The method further includes de-allocating the first proxy endpoint, when the first topology is deactivated making the first pluggable compute module and the NVMe storage device available for a second topology.


