Storage Node Reintegration via Importance and Reliability Thresholds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In a storage cluster with multiple nodes, automated reintegration of a failed storage node can destabilize the system, especially if frequent failures occur, as it may lead to frequent reintegration attempts and compromise dataset redundancy.

Innovation Solution

A cluster configuration control apparatus determines the importance and reliability of storage nodes, automatically reintegrating nodes only when their importance and reliability meet predetermined thresholds, ensuring stability by minimizing unnecessary reintegration and maintaining dataset accessibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automated reintegration is performed for storage nodes that fail, then the burden on administrators is reduced and node availability is improved, but the stability of the storage cluster deteriorates due to frequent reintegration attempts

Engineering Contradiction:
Improveautomated reintegrationVSAvoidstorage cluster stability
Core Design Contradiction:
Extent of automationVSStability of the object's composition

Solution Approach 1:

The system changes the parameter of reintegration decision-making from automatic to conditional based on importance and reliability assessments. By evaluating multiple parameters (node importance, failure frequency, data redundancy) before allowing reintegration, the system prevents automated reintegration of unstable nodes while maintaining automation for stable nodes.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system implements feedback mechanisms by monitoring node behavior, failure patterns, and cluster stability. The reintegration decision is based on feedback from reliability determination (failure frequency monitoring) and importance determination (data redundancy assessment), allowing the system to adapt reintegration policies based on observed node performance.

Inventive Principle:
Principle #23Feedback

2Reliability

If storage nodes are automatically reintegrated after failure, then node availability is improved, but dataset redundancy deteriorates due to frequent failures and reintegrations

Engineering Contradiction:
Improvenode availabilityVSAvoiddataset redundancy
Core Design Contradiction:
ReliabilityVSLoss of substance

Solution Approach 1:

The system performs preliminary assessment of node importance and reliability before allowing reintegration. By evaluating whether data redundancy requirements can be met and whether the node is stable enough for reintegration, the system prevents reintegration actions that would compromise dataset redundancy or cluster stability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes the reintegration parameter from unconditional to conditional based on importance determination. By assessing data redundancy levels and node criticality, the system selectively allows reintegration only when it will not compromise dataset redundancy, thus preserving substance while maintaining availability.

Inventive Principle:
Principle #35Parameter changes

3Stability of the object's composition

If manual reintegration is performed by administrators, then storage cluster stability is maintained through careful assessment, but the burden on administrators increases

Engineering Contradiction:
Improvestorage cluster stabilityVSAvoidadministrator burden
Core Design Contradiction:
Stability of the object's compositionVSEase of operation

Solution Approach 1:

The system enables self-service by automatically performing reintegration for nodes that meet stability and redundancy criteria. The node management apparatus autonomously evaluates importance and reliability, and executes reintegration without administrator intervention for suitable cases, reducing operational burden while maintaining stability through systematic assessment.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system applies partial automation by reintegrating only those nodes that pass the importance and reliability thresholds, rather than automatically reintegrating all failed nodes. This partial action approach maintains stability for critical nodes while automating reintegration for stable, non-critical nodes, balancing administrator burden reduction with cluster stability.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10795587B2Storage system and cluster configuration control method
Publication Date: 2020.10.06 HITACHI VANTARA LTD
  • US10795587B2 patent drawing
  • US10795587B2 patent drawing
  • US10795587B2 patent drawing

AI summary

It is determined whether the importance of an object storage node is equal to or larger than a predetermined importance and the reliability of the object storage node is equal to or larger than a predetermined reliability, the object storage node being a storage node set as an object among N storage nodes that are members of a storage cluster, N being an integer equal to or larger than 3. When the determination result is true, reintegration of the object storage node is performed. The importance of the object storage node depends on highness of availability when assuming that the object storage node has left the storage cluster. The reliability of the object storage node depends on the tendency of operation of the object storage node.