Storage Node Failover via Localized Fault Domain Rebalancing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Software Defined Storage (SDS) systems face significant rebalance times and vulnerability to data loss due to storage node failures, especially in large-scale environments like those required for emerging AI workloads, where traditional cluster-wide rebalancing is time-consuming and risky.
Innovation Solution
The solution involves disaggregating storage nodes and using low-latency, high-bandwidth interfaces like CXL, along with storage software innovations to distribute data across fault domains, allowing for instantaneous failover and data replication via PCIe non-transparent bridging and unordered stream writes, thereby avoiding cluster-wide rebalancing and reducing data loss risks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional cluster-wide rebalancing is used to maintain resiliency after storage node failure, then data resiliency goals are maintained, but rebalance time increases significantly (several hours) and cluster vulnerability to cascading failures increases
Solution Approach 1:
The patent segments the rebalancing process by limiting data migration to only the affected fault domain rather than performing cluster-wide rebalancing. The system identifies the specific fault domain containing the failed storage node and restricts rebalancing operations to that localized area, thereby reducing overall rebalance time while maintaining resiliency goals.
Solution Approach 2:
The patent applies local quality by making rebalancing decisions based on local fault domain characteristics rather than global cluster state. The system evaluates and performs rebalancing operations independently within each fault domain, allowing faster localized recovery without triggering cluster-wide operations.
2Quantity of substance
If storage node capacity is increased to meet AI workload demands, then storage capacity is improved, but rebalance time increases and service level agreement impacts increase
Solution Approach 1:
The patent segments the storage cluster into fault domains that can be rebalanced independently. When a storage node fails in a multi-PB capacity environment, only the affected fault domain undergoes rebalancing rather than the entire cluster, significantly reducing rebalance time even as storage capacity increases.
Solution Approach 2:
The patent implements dynamic fault domain identification and isolation mechanisms that adapt to the specific failure conditions. The system dynamically determines which fault domains require rebalancing based on the failure location and data distribution, enabling flexible response times that scale with storage capacity requirements.
3Reliability
If cluster-wide rebalancing is performed to maintain data distribution, then data resiliency is maintained, but the cluster becomes vulnerable to cascading failures during the rebalancing process
Solution Approach 1:
The patent segments the rebalancing operation to affect only specific fault domains rather than the entire cluster. By isolating rebalancing to localized areas, the system reduces the footprint of potential failure propagation during the rebalancing process, thereby minimizing cascading failure risks while maintaining data resiliency.
Solution Approach 2:
The patent implements beforehand cushioning by pre-configuring fault domain isolation mechanisms and pre-positioning data replicas within localized fault domains. This preparation creates a buffer that prevents failure propagation during rebalancing operations, protecting the cluster from cascading failures before they can occur.
Data Source
AI summary
Systems, apparatuses and methods may provide for technology that detects a first failure in a first storage server, wherein the first storage server is connected to a first non-volatile memory (NVM) via a switch, selects a second storage server that is connected to the first NVM via the switch, wherein the first storage server and the second storage server are in a storage cluster, and configures the second storage server to host first data resident on the first NVM, wherein configuring the second storage server to host the first data bypasses a cluster-wide rebalance of the storage cluster.


