Affinity-Based Data Distribution for Cluster Storage Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage systems face inefficiencies in data recovery due to the concentration of protection groups among a small subset of storage entities, leading to slowed recovery processes when a node or disk fails, as only a limited set of entities are heavily involved in the recovery operation.
Innovation Solution
The implementation of affinity-based data distribution logic, which tracks and manages the affinity levels between storage entities to distribute protection groups evenly, ensuring that multiple storage entities participate in recovery operations, thereby reducing the load on specific nodes or disks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If capacity load balancing techniques are used to distribute data evenly across storage entities, then data distribution uniformity is improved, but protection groups become concentrated among small subsets of storage entities, worsening recovery speed
Solution Approach 1:
The system segments protection groups into different subsets and assigns each subset to different storage entities, ensuring that no single storage entity becomes a bottleneck for recovery operations. This segmentation allows multiple storage entities to participate in parallel recovery processes.
Solution Approach 2:
The patent introduces a new dimension to data distribution by tracking affinity levels between storage entities and using this affinity information to make distribution decisions. This transforms the traditional single-dimension load balancing into a multi-dimensional distribution strategy that considers both load balance and recovery performance.
2Device complexity
If protection groups are concentrated among small subsets of storage entities for simplified management, then system complexity is reduced, but recovery operations are slowed due to limited participant entities
Solution Approach 1:
The system implements feedback mechanisms by tracking affinity levels between storage entities and using this information to dynamically adjust protection group distribution. This feedback loop enables the system to automatically optimize for both management simplicity and recovery performance without manual intervention.
Solution Approach 2:
The patent performs preliminary distribution of protection groups based on predicted affinity patterns before failures occur. By pre-distributing groups across multiple storage entities with low affinity relationships, the system prepares the infrastructure for efficient parallel recovery operations without adding complexity to the distribution management.
3Productivity
If affinity-based data distribution logic is implemented to distribute protection groups evenly, then recovery efficiency is improved, but system complexity increases due to affinity tracking and management
Solution Approach 1:
The system employs dynamic affinity tracking that adapts to changing storage entity relationships over time. Rather than using static distribution rules, the affinity-based logic continuously monitors and adjusts distribution decisions based on current system state, enabling efficient recovery while managing complexity through adaptability.
Solution Approach 2:
The patent changes the distribution parameter from simple load balance metrics to affinity-based metrics. By using affinity levels as the key parameter for distribution decisions, the system achieves better recovery efficiency while the affinity tracking infrastructure manages the complexity through a unified parameter framework.
Data Source
AI summary
The described technology is generally directed towards distributing data fragments and coding fragments of a protection group among storage entities (e.g., nodes or disks) based on affinity levels (e.g., maintained in an affinity matrix) that represent dependency relationships between the storage entities with respect to storing protection groups. The technology operates to distribute a protection group's components such that the affinity level between any pair of storage entities is approximately the same as any other pair. In the event of a storage entity failure, as a result of the affinity-based distribution of the protection group components needed for data recovery, a larger number of the other storage entities can be involved in the data recovery (relative to the number likely involved without affinity-based distribution). This tends to assure a better load balance and faster data recovery.


