Storage Node Failover via Localized Fault Domain Rebalancing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Software Defined Storage (SDS) systems face significant rebalance times and vulnerability to data loss due to storage node failures, especially in large-scale environments like those required for emerging AI workloads, where traditional cluster-wide rebalancing is time-consuming and risky.

Innovation Solution

The solution involves disaggregating storage nodes and using low-latency, high-bandwidth interfaces like CXL, along with storage software innovations to distribute data across fault domains, allowing for instantaneous failover and data replication via PCIe non-transparent bridging and unordered stream writes, thereby avoiding cluster-wide rebalancing and reducing data loss risks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional cluster-wide rebalancing is used to maintain resiliency after storage node failure, then data resiliency goals are maintained, but rebalance time increases significantly (several hours) and cluster vulnerability to cascading failures increases

Engineering Contradiction:
Improvedata resiliencyVSAvoidrebalance time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the rebalancing process by limiting data migration to only the affected fault domain rather than performing cluster-wide rebalancing. The system identifies the specific fault domain containing the failed storage node and restricts rebalancing operations to that localized area, thereby reducing overall rebalance time while maintaining resiliency goals.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by making rebalancing decisions based on local fault domain characteristics rather than global cluster state. The system evaluates and performs rebalancing operations independently within each fault domain, allowing faster localized recovery without triggering cluster-wide operations.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If storage node capacity is increased to meet AI workload demands, then storage capacity is improved, but rebalance time increases and service level agreement impacts increase

Engineering Contradiction:
Improvestorage capacityVSAvoidrebalance time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent segments the storage cluster into fault domains that can be rebalanced independently. When a storage node fails in a multi-PB capacity environment, only the affected fault domain undergoes rebalancing rather than the entire cluster, significantly reducing rebalance time even as storage capacity increases.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic fault domain identification and isolation mechanisms that adapt to the specific failure conditions. The system dynamically determines which fault domains require rebalancing based on the failure location and data distribution, enabling flexible response times that scale with storage capacity requirements.

Inventive Principle:
Principle #15Dynamics

3Reliability

If cluster-wide rebalancing is performed to maintain data distribution, then data resiliency is maintained, but the cluster becomes vulnerable to cascading failures during the rebalancing process

Engineering Contradiction:
Improvedata resiliencyVSAvoidcascading failures
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent segments the rebalancing operation to affect only specific fault domains rather than the entire cluster. By isolating rebalancing to localized areas, the system reduces the footprint of potential failure propagation during the rebalancing process, thereby minimizing cascading failure risks while maintaining data resiliency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements beforehand cushioning by pre-configuring fault domain isolation mechanisms and pre-positioning data replicas within localized fault domains. This preparation creates a buffer that prevents failure propagation during rebalancing operations, protecting the cluster from cascading failures before they can occur.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS20230013798A1Cluster wide rebuild reduction against storage node failures
Publication Date: 2023.01.19 INTEL CORP
  • US20230013798A1 patent drawing
  • US20230013798A1 patent drawing
  • US20230013798A1 patent drawing

AI summary

Systems, apparatuses and methods may provide for technology that detects a first failure in a first storage server, wherein the first storage server is connected to a first non-volatile memory (NVM) via a switch, selects a second storage server that is connected to the first NVM via the switch, wherein the first storage server and the second storage server are in a storage cluster, and configures the second storage server to host first data resident on the first NVM, wherein configuring the second storage server to host the first data bypasses a cluster-wide rebalance of the storage cluster.