Dynamic Fault Tolerance Policy for Object Storage Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing techniques for managing fault tolerance policies in distributed object-based storage systems, such as vSAN, are limited by manual reconfiguration requirements that do not scale well in larger data centers and cannot be applied to all objects, including those not exposed to users like metadata objects.

Innovation Solution

Implementing dynamic fault tolerance policies that automatically adjust the number of host failures to tolerate (HFT) based on real-time conditions in a storage cluster, allowing for optimal fault tolerance configurations without manual intervention, even for objects not exposed to users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If manual reconfiguration of fault tolerance policies is performed, then policy changes can be made for exposed objects, but the process does not scale well in larger data centers and cannot be applied to non-exposed objects like metadata objects

Engineering Contradiction:
Improvepolicy reconfiguration capabilityVSAvoidoperational complexity
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements dynamic fault tolerance policies that automatically adjust the number of host failures to tolerate (HFT) based on real-time conditions in the storage cluster. This transforms the static, manual policy reconfiguration process into a dynamic, automated system that adapts to changing cluster conditions without requiring manual intervention for each policy change.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system enables self-service by allowing the storage cluster to automatically reconfigure fault tolerance policies based on monitored conditions. The patent describes how the system can autonomously determine when policy changes are needed and implement them without user intervention, including for non-exposed objects like metadata objects that users cannot directly access or modify.

Inventive Principle:
Principle #25Self-service

2Reliability

If the number of host failures to tolerate (HFT) is increased for better protection, then fault tolerance improves, but storage resources are consumed more rapidly

Engineering Contradiction:
Improvefault toleranceVSAvoidstorage capacity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent implements dynamic adjustment of HFT values based on real-time monitoring of storage cluster conditions. The system automatically increases or decreases the number of host failures to tolerate based on available resources and current cluster state, allowing optimal fault tolerance protection while efficiently utilizing storage capacity without manual intervention.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system dynamically changes the HFT parameter based on cluster conditions. The patent describes how the number of host failures to tolerate is adjusted as a variable parameter rather than a fixed value, allowing the system to optimize between reliability and resource consumption by modifying this key parameter in response to changing conditions.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11693559B2Dynamic object policy reconfiguration mechanism for object storage system
Publication Date: 2023.07.04 VMWARE INC
  • US11693559B2 patent drawing
  • US11693559B2 patent drawing
  • US11693559B2 patent drawing

AI summary

A method for dynamic storage object configuration in a datacenter is provided. Embodiments include determining a number of fault domains in a storage cluster that have sufficient storage capacity for creating a storage object. Embodiments include applying a dynamic fault tolerance policy to the number of fault domains that have sufficient capacity for creating the storage object in order to determine a number of host failures to tolerate for the storage object, the dynamic fault tolerance policy specifying a manner of determining, for any respective storage object, a respective number of host failures to tolerate for storing the respective storage object in a respective storage cluster based on at least a respective number of fault domains of the respective storage cluster. Embodiments include implementing the storage object on the storage cluster based on the number of host failures to tolerate for the storage object.