Virtual Switch Endpoints for Disaggregated Accelerator RDMA

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In data centers with disaggregated resources, flow management up to the endpoint (e.g., accelerator devices) is not possible, and endpoint services are not exposed to an orchestration layer, limiting efficient resource allocation and utilization.

Innovation Solution

Implementing a system that manages disaggregated accelerator networks by configuring virtual network connections between compute sleds and accelerator sleds, allowing for remote direct memory access (RDMA) and exposing accelerator devices as switchable endpoints to an upper orchestration layer, enabling flow management and network efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If disaggregated accelerator devices are used in data centers, then resource flexibility and scalability are improved, but flow management capability and endpoint service exposure deteriorate

Engineering Contradiction:
Improveresource flexibilityVSAvoidflow management capability
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent introduces virtual switch endpoints as intermediary objects that bridge the gap between disaggregated accelerator devices and the orchestration layer. These virtual endpoints expose accelerator services to the orchestration layer without requiring direct access to physical devices, enabling flow management while preserving resource flexibility. The virtual switch endpoints act as mediators that translate orchestration layer commands into device-specific operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system segments the network architecture into distinct layers: the orchestration layer, the virtualization layer with virtual switch endpoints, and the physical device layer with disaggregated accelerators. This segmentation allows each layer to operate independently, maintaining resource flexibility at the device level while enabling centralized flow management at the orchestration layer through virtual endpoint abstractions.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If disaggregated accelerator devices are used, then resource allocation flexibility is improved, but network efficiency and resource utilization deteriorate

Engineering Contradiction:
Improveresource allocation flexibilityVSAvoidnetwork efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent implements feedback mechanisms where the orchestration layer monitors resource utilization and performance metrics of disaggregated accelerators, then dynamically adjusts resource allocation and flow management policies. This feedback loop enables the system to optimize network efficiency while maintaining allocation flexibility, as the orchestration layer can respond to actual device performance and workload demands in real-time.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system employs dynamic resource allocation where virtual switch endpoints and network flows can be dynamically created, modified, and deleted based on workload requirements. The orchestration layer can dynamically assign different accelerators to different workloads, adjust bandwidth allocation, and reconfigure network paths, thereby maintaining both allocation flexibility and network efficiency through adaptive resource management.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11228539B2Technologies for managing disaggregated accelerator networks based on remote direct memory access
Publication Date: 2022.01.18 INTEL CORP
  • US11228539B2 patent drawing
  • US11228539B2 patent drawing
  • US11228539B2 patent drawing

AI summary

Technologies for network interface controllers (NICs) include a compute sled and an accelerator sled in communication over a network. The accelerator sled configures a virtual switch endpoint associated with a remote direct memory access (RDMA) server instance that is associated with a field-programmable gate array (FPGA) of the accelerator sled. The accelerator sled updates local software defined networking (SDN) tables with a virtual tunnel associated with the virtual switch endpoint and a remote compute sled. A virtual switch of the accelerator sled switches virtual tunnel traffic from the remote compute sled to the RDMA server instance, which transfers data to or from the FPGA. The compute sled also updates a local SDN table with the virtual tunnel, and a virtual switch of the compute sled switches virtual tunnel traffic to or from the accelerator sled. Other embodiments are described and claimed.