Distributed Storage Redundancy Recovery via Segment Migration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The 'federated array of bricks' (FAB) architecture for distributed data-storage systems faces challenges in efficiently managing and recovering redundancy under various failure conditions, particularly in managing and recovering data after the failure of individual mass-storage devices within component data-storage systems.
Innovation Solution
The method involves migrating affected segments from the failed component data-storage system to other systems within the distributed data-storage system, ensuring only sufficient segments are moved to provide necessary free space for redundancy recovery, leveraging redundancy-recovery operations such as data mirroring and erasure coding to restore data integrity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If segments are migrated to other component data-storage systems upon mass-storage device failure, then redundancy is recovered and data integrity is restored, but data movement overhead and system complexity increase
Solution Approach 1:
The patent divides the distributed data-storage system into discrete component data-storage systems (bricks), each independently manageable. When a mass-storage device fails, only the affected segments are identified and migrated, rather than requiring system-wide redundancy recovery operations. This segmentation enables localized recovery actions that reduce overall system complexity while maintaining data integrity.
Solution Approach 2:
The system pre-establishes a pool of free space across component data-storage systems before failures occur. Upon detecting mass-storage device failure, the system immediately utilizes this pre-positioned free space to receive migrated segments, eliminating the need for complex real-time space allocation and coordination during recovery operations.
2Productivity
If sufficient free space is maintained in component data-storage systems for redundancy recovery, then recovery operations can proceed efficiently, but storage capacity available for user data is reduced
Solution Approach 1:
The patent implements a free-space pool mechanism where component data-storage systems maintain a predetermined amount of free space dedicated to redundancy recovery operations. This partial allocation ensures that sufficient capacity is always available for efficient recovery without requiring excessive free space that would unnecessarily reduce user-available storage capacity.
Solution Approach 2:
When mass-storage devices fail, the system migrates affected segments to component data-storage systems with available free space in the pool. This copying approach allows the system to restore redundancy by creating duplicate copies of critical data segments in locations with pre-allocated free space, enabling efficient recovery without permanently sacrificing user storage capacity.
3Loss of energy
If only sufficient segments are migrated to provide necessary free space, then data movement is minimized and recovery is optimized, but the process requires precise calculation and coordination
Solution Approach 1:
The system continuously monitors free space availability across the pool of component data-storage systems and adjusts segment migration decisions based on real-time feedback. When a mass-storage device fails, the system calculates the precise amount of free space needed for recovery and migrates only that sufficient number of segments, avoiding unnecessary data movement while maintaining optimal recovery coordination.
Data Source
AI summary
Embodiments of the present invention are directed to methods, and distributed data-storage systems employing the methods, for recovering redundancy within a distributed data-storage system upon failure of one or more mass-storage devices within a component data-storage system of the distributed data-storage system. In certain embodiments, failure of a mass-storage device within a component data-storage system elicits a redundancy-recovery operation in which segments affected by the mass-storage-device failure or failures are moved, by a process referred to as “migration,” to other component data-storage systems of the distributed data-storage system, and are recovered as a by-product of migration. Certain embodiments of the present invention more efficiently address redundancy recovery by moving only as many segments from the component data-storage system as needed to provide sufficient free space within the component data-storage system to recover the remaining segments affected by the mass-storage-device failure or failures within the component data-storage system.


