Cluster Configuration Replication via Selective Schema

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In cluster storage environments, issues arise when information is not reliably replicated between nodes and storage clusters, leading to inaccessible user data and inefficient failover operations during outages, due to lack of efficient replication and detection mechanisms.

Innovation Solution

Implementing a cluster configuration schema to selectively replicate storage operations, deploying cluster-wide service agents with a master and standby configuration, and defining outage detection metrics to ensure high availability and swift recovery.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If all cluster configuration information is replicated between nodes, then reliability of data access is improved, but resource usage and system complexity increase

Engineering Contradiction:
Improvereliability of data accessVSAvoidresource usage
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent segments cluster configuration information into different types (service-level vs. cluster-level) and applies selective replication only to service-level information. This segmentation allows the system to maintain reliability for critical services while reducing overall resource consumption by not replicating all configuration data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality by applying different replication strategies to different types of configuration information based on their specific requirements. Service-level information receives selective replication to maintain local availability, while cluster-level information uses different handling, optimizing resource usage for each information type.

Inventive Principle:
Principle #3Local quality

2Reliability

If comprehensive replication of cluster information is implemented, then failover reliability is improved, but replication time and system complexity increase

Engineering Contradiction:
Improvefailover reliabilityVSAvoidreplication time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent divides configuration information into service-level and cluster-level categories, replicating only service-level information during failover operations. This segmentation reduces replication time by excluding unnecessary cluster-level information while maintaining sufficient reliability for service continuity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by replicating only the necessary subset of configuration information (service-level) required for failover, rather than performing comprehensive replication of all cluster information. This reduces replication time while providing adequate failover capability.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If cluster-wide service agents are deployed on all nodes, then service availability is improved, but device complexity and resource usage increase

Engineering Contradiction:
Improveservice availabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges the functionality of multiple service agents into a single cluster-wide service agent that operates on one node. This consolidation maintains service availability by centralizing service management while reducing device complexity by eliminating redundant agent installations on other nodes.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The cluster-wide service agent is designed with universal functionality to handle service management across the entire cluster from a single node. This multi-functional agent replaces the need for separate agents on each node, reducing complexity while maintaining comprehensive service coverage.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Loss of energy

If selective replication based on cluster configuration schema is used, then resource usage is reduced, but measurement precision of replication control increases

Engineering Contradiction:
Improveresource usageVSAvoidprecision of replication control
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The patent uses the cluster configuration schema to segment information into replicable and non-replicable categories. This precise segmentation reduces resource usage by replicating only necessary information while the schema provides the measurement precision needed to correctly identify which information should be replicated.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The cluster configuration schema acts as an intermediary that mediates between the need for selective replication and the requirement for precise control. It provides the structured framework that enables resource-efficient replication while maintaining accurate control over what information is replicated.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11226777B2Cluster configuration information replication
Publication Date: 2022.01.18 NETAPP INC
  • US11226777B2 patent drawing
  • US11226777B2 patent drawing
  • US11226777B2 patent drawing

AI summary

One or more techniques and/or systems are provided for cluster configuration information replication, managing cluster-wide service agents, and/or for cluster-wide outage detection. In an example of cluster configuration information replication, a replication workflow corresponding to a storage operation implemented for a storage object (e.g., renaming of a volume) of a first cluster may be transferred to a second storage cluster for selectively implementation. In an example of managing cluster-wide service agents, cluster-wide service agents are deployed to nodes of a cluster storage environment, where a master agent actively processes cluster service calls and standby agents passively wait for reassignment as a failover master in the event the master agent fails. In an example of cluster-wide outage detection, a cluster-wide outage may be determined for a cluster storage environment based upon a number of inaccessible nodes satisfying a cluster outage detection metric.