Self-Healing IPU Networks for Host CPU Failure Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
When a host CPU or Infrastructure Processing Unit (IPU) fails in a server environment, the connected hardware accelerators, such as GPUs and TPUs, become unavailable, leading to poor resource utilization and workload migration issues, even if these accelerators are functional.
Innovation Solution
A self-healing network of Infrastructure Processing Units (IPUs) and associated methods that seamlessly migrate IPU responsibilities to another IPU in a data center network, using a central Resource Manager and foundational network interface controllers (fNICs) to manage XPU clusters, encapsulating PCIe transactions into network packets for remote communication and managing XPU failures through heartbeat signals and recovery attempts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If XPUs are connected to a host CPU on a server, then the XPUs can perform compute operations, but if the host CPU fails, the entire server becomes unavailable and workloads must be migrated to another server
Solution Approach 1:
The system separates the control plane (IPU) from the data plane (XPU), allowing the XPU to remain available even if its local host CPU fails. The IPU manages multiple XPUs independently, so failure of one host CPU does not affect other XPUs managed by the IPU.
Solution Approach 2:
The IPU acts as an intermediary between the host CPU and XPUs, providing a layer of indirection that allows XPUs to be managed and accessed even when the local host CPU fails. The IPU can redirect workloads to different host CPUs or maintain XPU availability through its own processing capabilities.
2Productivity
If control and management is moved to the IPU with no local host CPU, then the host CPU is freed for compute operations, but if the IPU fails, the XPUs become unavailable even though resources are healthy
Solution Approach 1:
The system implements preliminary failure detection mechanisms where the IPU monitors the health of host CPUs and XPUs. When failures are detected, the IPU has pre-established pathways and procedures to rapidly redirect workloads, minimizing downtime and maintaining XPU availability.
Solution Approach 2:
The system implements redundancy and failover capabilities where backup host CPUs or alternative IPU instances are prepared in advance. When an IPU fails, the system can activate standby resources to maintain XPU availability, cushioning against the failure impact.
3Productivity
If the IPU manages XPU clusters without local host CPU, then resource utilization should improve, but failure detection and recovery complexity increases
Solution Approach 1:
The system implements comprehensive feedback mechanisms where XPUs continuously report their status to the IPU, and the IPU monitors host CPU health. This feedback loop enables automatic detection of failures and triggers appropriate recovery actions, simplifying the management of complex IPU-based systems through automated responses.
Data Source
AI summary
Self-healing networks of Infrastructure Processing Units (IPUs) and associated methods and apparatus. The self-healing IPUs manage other processing units (XPU) clusters by seamlessly migrating the IPU responsibilities to another IPU in the networked environment (e.g., data center) that may be available when an IPU failures or becomes unavailable. A central Resource Manager is used to monitors the health of the IPUs in the data center and in the event of in IPU failure, locates another IPU and assigns it to take over the failed IPU's functions. Replacement and workload migration of a failed XPU in an IPU managed XPU cluster with a remote XPU that is network connected is also supported. The IPU monitors the health of the XPUs in its cluster an informs the Resource Manager of an XPU failure which locates another XPU in the data center and assigns it to the cluster that has the failed XPU.


