Automated Fault Tolerance for Hyper-Converged Storage Nodes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional hyper-converged systems lack the capability to track the physical layout of hardware storage nodes, leading to manual and labor-intensive configuration and auditing of fault tolerance, which becomes complex as deployments grow and systems are upgraded or replaced.

Innovation Solution

The system automatically detects and configures mirror hardware storage nodes to ensure they do not share physical equipment, using rack and chassis identifiers to establish and audit pools of data storage that meet target fault tolerance levels, thereby eliminating single points of failure.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual configuration and auditing of fault tolerance is used in hyper-converged systems, then system reliability can be maintained, but operational complexity and labor intensity increase significantly as deployments grow

Engineering Contradiction:
Improvefault toleranceVSAvoidconfiguration complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system automatically discovers the physical layout of storage nodes and performs fault tolerance configuration without manual intervention. The management software queries physical attributes (rack, chassis, slot identifiers) from storage nodes and autonomously determines mirror pairings based on fault tolerance requirements, eliminating the need for operators to manually track physical locations and configure mirroring.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system continuously audits the physical layout and mirror pairings to verify fault tolerance compliance. By querying physical attributes and comparing them against configured mirror pairs, the system provides feedback on whether the current configuration meets fault tolerance requirements, automatically identifying and reporting violations.

Inventive Principle:
Principle #23Feedback

2Productivity

If physical layout tracking is implemented to enable automated fault tolerance configuration, then operational efficiency improves, but system complexity increases due to additional tracking requirements

Engineering Contradiction:
Improveconfiguration efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The management software acts as an intermediary between storage nodes and operators, automatically querying physical attributes from storage nodes and translating them into fault tolerance configurations. This intermediary layer handles the complexity of physical layout tracking and mirror pairing logic, presenting a simplified interface to operators while maintaining comprehensive control over fault tolerance configuration.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If mirror hardware storage nodes are configured to not share physical equipment, then fault tolerance is improved, but storage capacity utilization decreases due to redundant mirroring

Engineering Contradiction:
Improvefault tolerance levelVSAvoidstorage capacity utilization
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system applies different physical separation requirements to different levels of the storage hierarchy. While mirror pairs must be physically separated at the rack or chassis level to ensure fault tolerance, the system optimally places these mirrors within the available storage capacity by selecting appropriate candidate nodes that meet the physical separation criteria while maximizing space utilization.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10785294B1Methods, systems, and computer readable mediums for managing fault tolerance of hardware storage nodes
Publication Date: 2020.09.22 EMC IP HLDG CO LLC
  • US10785294B1 patent drawing
  • US10785294B1 patent drawing
  • US10785294B1 patent drawing

AI summary

Methods, systems, and computer readable mediums for managing fault tolerance. A method includes receiving a request to establish a pool of data storage for an application of a distributed computing system. The distributed computing system includes hardware storage nodes integrated with compute nodes. The method includes receiving a target level of fault tolerance for the pool of data storage. The method includes establishing the pool of data storage by specifying, for each hardware storage node, a mirror hardware storage node for mirroring data stored on the hardware storage node so that the hardware storage node and the mirror hardware storage node do not share one or more pieces of physical equipment as specified in a physical layout of the hardware storage nodes to meet the target level of fault tolerance.