Interconnect Switch for Dynamic Accelerator Resource Sharing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data center systems face challenges in efficiently sharing accelerator resources among multiple node resources due to tight coupling in coherency domains and fault isolation requirements, leading to underutilization of resources and overprovisioning.
Innovation Solution
The implementation of an interconnect switch that allows for the dynamic reconfiguration of connections between node resources and accelerator resources, enabling hot-removal and hot-add operations without physical intervention, using a resource manager to manage these connections.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If accelerator resources are tightly coupled to node resources within coherency domains, then fault isolation is maintained, but resource utilization efficiency deteriorates due to inability to share accelerators across nodes
Solution Approach 1:
The system segments the connection between accelerators and nodes by introducing a switching fabric that decouples the traditional direct coupling. Accelerators are connected to the switching fabric independently from nodes, allowing the connection to be segmented and reconfigured dynamically while maintaining fault isolation through controlled connection paths.
Solution Approach 2:
A switching fabric acts as an intermediary between accelerators and nodes, enabling indirect connections that maintain fault isolation while allowing flexible resource sharing. The switch fabric mediates connections, permitting accelerators to be shared across multiple nodes without direct coupling, thus resolving the contradiction between isolation and utilization.
2Adaptability or versatility
If accelerator resources are physically moved between node resources, then resource allocation flexibility is improved, but system complexity and operational difficulty worsen due to physical intervention requirements
Solution Approach 1:
The system implements dynamic reconfiguration of accelerator connections through a programmable switching fabric. Connection paths can be changed in real-time through software control without physical movement of devices, enabling flexible resource allocation while maintaining ease of operation through automated management.
Solution Approach 2:
The patent replaces the mechanical/physical system of moving accelerators with an electrical/software-based switching system. The switching fabric electronically reconfigures connections through control signals, eliminating the need for physical intervention and reducing operational complexity while maintaining allocation flexibility.
3Reliability
If data centers are overprovisioned with accelerator resources to meet worst-case scenarios, then reliability and service level agreement compliance are improved, but resource underutilization worsens during normal operation
Solution Approach 1:
The switching fabric enables accelerators to serve multiple nodes and multiple functions dynamically. A single accelerator can be allocated to different nodes based on workload demands, allowing the system to meet peak demands with fewer total accelerators while maintaining high utilization during normal operation, thus reducing energy waste from underutilized resources.
Solution Approach 2:
The system dynamically changes connection parameters through the switching fabric, allowing accelerator allocation to adapt to varying workload conditions. During peak loads, accelerators are allocated to meet service level agreements; during normal operation, accelerators are consolidated to improve utilization and reduce energy consumption, effectively managing the trade-off between reliability and efficiency.
Data Source
AI summary
The present disclosure describes a number of embodiments related to devices and techniques for implementing an interconnect switch to provide a switchable low-latency bypass between node resources such as CPUs and accelerator resources for caching. A resource manager may be used to receive an indication of a node of a plurality of nodes and an indication of an accelerator resource of a plurality of accelerator resources to connect to the node. If the indicated accelerator resource is connected to another node of the plurality of nodes, then transmit, to a interconnect switch, one or more hot-remove commands. The resource manager may then transmit to the interconnect switch one or more hot-add commands to connect the node resource and the accelerator resource.


