Cluster Data Redundancy via Generation Identifiers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cluster architectures for data backup and recovery may lead to significant data loss if a portion of the data is not adequately protected, as they spread data across multiple members to minimize loss in case of failure, which can be intolerable for clients depending on the data.
Innovation Solution
The solution involves a cluster environment with multiple member indexers that use data replication and generation identifiers (GEN_IDs) to manage data storage and recovery, where a forwarder device selects a primary indexer and specifies a replication factor, and secondary indexers are designated based on availability, with transitions in primacy managed using GEN_IDs to ensure data accessibility and redundancy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is spread across multiple cluster members to minimize loss, then data loss is reduced, but data availability and完整性 cannot be guaranteed for significant data portions
Solution Approach 1:
The patent segments data into distinct buckets that can be independently managed and replicated. Each bucket represents a logical unit of data that can be assigned to specific cluster members, allowing selective replication strategies for different data portions based on their importance and access patterns.
Solution Approach 2:
The patent applies different replication factors to different data buckets based on their specific requirements. Critical data buckets can have higher replication factors stored across multiple members, while less critical data can have lower replication factors, optimizing both data protection and resource utilization.
2Reliability
If data is replicated across multiple indexers, then data availability is improved, but system complexity increases
Solution Approach 1:
The patent implements self-healing capabilities where the cluster automatically detects failures, reassignes data buckets to available members, and maintains replication integrity without manual intervention. The system monitors its own state and performs corrective actions autonomously.
Solution Approach 2:
The patent employs dynamic replication factor adjustment based on real-time cluster conditions. Replication factors can be modified according to data importance, available resources, and failure scenarios, allowing the system to adapt its complexity level to current operational needs.
3Reliability
If replication factor is increased for critical data, then data protection is enhanced, but storage resources are consumed
Solution Approach 1:
The patent applies differentiated replication strategies to different data buckets based on their criticality. High-priority data receives higher replication factors with multiple copies across the cluster, while lower-priority data uses minimal replication, optimizing the balance between protection and resource consumption.
Solution Approach 2:
The patent dynamically adjusts replication factors as a configurable parameter based on data importance, failure probability, and resource availability. This allows flexible optimization of storage resource allocation while maintaining adequate protection levels for critical data.
Data Source
AI summary
Embodiments are directed towards managing within a cluster environment having a plurality of indexers for data storage using redundancy the data being managed using a generation identifier, such that a primary indexer is designated for a given generation of data. When a master device for the cluster fails, data may continue to be stored using redundancy, and data searches performed may still be performed.


