Node Capability Assessment in Storage Cluster High Availability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for determining whether a node in a storage cluster system has sufficient resources to take over the functions of another node in case of failure are oversimplified, leading to inaccurate and unusable results.

Innovation Solution

An administration system that monitors node usage levels and performance, generates node models correlating usage types with performance metrics, and analyzes these models to determine if nodes within a high-availability group can effectively take over each other's functions, providing graphical diagnostics and alerts for resource thresholds.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If simplified approaches are used to determine node headroom, then the assessment process is fast and simple, but the results become inaccurate and unusable

Engineering Contradiction:
Improvesimplicity of assessment processVSAvoidaccuracy of node capability assessment
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent transforms the assessment from a static simplified check to a dynamic multi-parameter analysis. It introduces comprehensive parameters including current resource utilization levels, historical performance data, workload characteristics, and predicted future states. The system calculates headroom based on multiple dimensions (CPU, memory, I/O, network) rather than a single oversimplified metric, resolving the contradiction by making the process sufficiently complex to be accurate but automated to remain operable.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements continuous monitoring and feedback loops where actual node performance and resource consumption are tracked over time. This feedback is used to refine headroom calculations and update node capability assessments dynamically. The system compares predicted versus actual performance and adjusts its models accordingly, ensuring high accuracy while maintaining automated operation through iterative refinement.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If comprehensive resource monitoring is implemented to accurately assess node headroom, then assessment accuracy improves, but system complexity and resource overhead increase

Engineering Contradiction:
Improveaccuracy of node capability assessmentVSAvoidcomplexity of monitoring system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a multi-functional assessment system that simultaneously performs multiple tasks: monitoring current resource usage, analyzing historical performance patterns, predicting future resource needs, calculating headroom for failover scenarios, and generating alerts. By consolidating these functions into a unified platform, the system achieves comprehensive accuracy without proportionally increasing complexity, as the same infrastructure serves multiple purposes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent performs preliminary analysis by continuously collecting and pre-processing resource utilization data before failover events occur. It establishes baseline performance metrics and identifies trends in advance, so that when headroom assessment is needed, the system can quickly query pre-computed data rather than gathering everything from scratch. This reduces the computational burden at decision points while maintaining comprehensive monitoring accuracy.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If real-time analysis of node models is performed to determine failover capability, then the system can prevent failures proactively, but processing time and computational resources increase

Engineering Contradiction:
Improvefault tolerance of storage clusterVSAvoidprocessing time for capability assessment
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary computations by continuously building and updating node models with historical performance data in the background. It pre-calculates performance baselines, identifies workload patterns, and maintains ready-to-use capability assessments. When a failover event is detected or anticipated, the system queries these pre-computed models rather than performing full analysis in real-time, significantly reducing decision latency while maintaining proactive failure prevention capability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic assessment where the system adapts its monitoring and analysis intensity based on system conditions. During normal operation, it uses lighter-weight continuous monitoring with pre-computed models. When anomalies are detected or failover risks increase, it dynamically intensifies analysis depth and frequency. This dynamic approach maintains high reliability for critical decisions while reducing average processing time during stable periods.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10031822B2Techniques for estimating ability of nodes to support high availability functionality in a storage cluster system
Publication Date: 2018.07.24 NETAPP INC
  • US10031822B2 patent drawing
  • US10031822B2 patent drawing
  • US10031822B2 patent drawing

AI summary

Various embodiments are generally directed to techniques for determining whether one node of a HA group is able to take over for another. An apparatus includes a model derivation component to derive a model correlating node usage level to node data propagation latency through and to node resource utilization from a first model of a first node of a storage cluster system and a second model of a second node of the storage cluster system, the first model based on a first usage level of the first node under a first usage type, and the second model based on a second usage level of the second node under a second usage type; and an analysis component to determine whether the first node is able to take over for the second node based on applying to the derived model a total usage level derived from the first and second usage levels.