Storage Node Reintegration via Importance and Reliability Thresholds
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In a storage cluster with multiple nodes, automated reintegration of a failed storage node can destabilize the system, especially if frequent failures occur, as it may lead to frequent reintegration attempts and compromise dataset redundancy.
Innovation Solution
A cluster configuration control apparatus determines the importance and reliability of storage nodes, automatically reintegrating nodes only when their importance and reliability meet predetermined thresholds, ensuring stability by minimizing unnecessary reintegration and maintaining dataset accessibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If automated reintegration is performed for storage nodes that fail, then the burden on administrators is reduced and node availability is improved, but the stability of the storage cluster deteriorates due to frequent reintegration attempts
Solution Approach 1:
The system changes the parameter of reintegration decision-making from automatic to conditional based on importance and reliability assessments. By evaluating multiple parameters (node importance, failure frequency, data redundancy) before allowing reintegration, the system prevents automated reintegration of unstable nodes while maintaining automation for stable nodes.
Solution Approach 2:
The system implements feedback mechanisms by monitoring node behavior, failure patterns, and cluster stability. The reintegration decision is based on feedback from reliability determination (failure frequency monitoring) and importance determination (data redundancy assessment), allowing the system to adapt reintegration policies based on observed node performance.
2Reliability
If storage nodes are automatically reintegrated after failure, then node availability is improved, but dataset redundancy deteriorates due to frequent failures and reintegrations
Solution Approach 1:
The system performs preliminary assessment of node importance and reliability before allowing reintegration. By evaluating whether data redundancy requirements can be met and whether the node is stable enough for reintegration, the system prevents reintegration actions that would compromise dataset redundancy or cluster stability.
Solution Approach 2:
The system changes the reintegration parameter from unconditional to conditional based on importance determination. By assessing data redundancy levels and node criticality, the system selectively allows reintegration only when it will not compromise dataset redundancy, thus preserving substance while maintaining availability.
3Stability of the object's composition
If manual reintegration is performed by administrators, then storage cluster stability is maintained through careful assessment, but the burden on administrators increases
Solution Approach 1:
The system enables self-service by automatically performing reintegration for nodes that meet stability and redundancy criteria. The node management apparatus autonomously evaluates importance and reliability, and executes reintegration without administrator intervention for suitable cases, reducing operational burden while maintaining stability through systematic assessment.
Solution Approach 2:
The system applies partial automation by reintegrating only those nodes that pass the importance and reliability thresholds, rather than automatically reintegrating all failed nodes. This partial action approach maintains stability for critical nodes while automating reintegration for stable, non-critical nodes, balancing administrator burden reduction with cluster stability.
Data Source
AI summary
It is determined whether the importance of an object storage node is equal to or larger than a predetermined importance and the reliability of the object storage node is equal to or larger than a predetermined reliability, the object storage node being a storage node set as an object among N storage nodes that are members of a storage cluster, N being an integer equal to or larger than 3. When the determination result is true, reintegration of the object storage node is performed. The importance of the object storage node depends on highness of availability when assuming that the object storage node has left the storage cluster. The reliability of the object storage node depends on the tendency of operation of the object storage node.


