Stripe Map Data Re-striping for Storage Cluster Expansion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current storage systems face inefficiencies in expanding data storage capacity through horizontal capacity expansion, as the process of re-striping data when adding a new node is lengthy and disruptive, requiring massive data movement and resulting in decreased performance and potential capacity crunches before migration is complete.
Innovation Solution
The system generates and uses stripe maps to minimize data movement during node addition, allowing the new node to contribute capacity immediately, with data re-stripping that only moves necessary data to the new node and maintains existing node operations, reducing I/O operations and migration time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is re-striped across all nodes when a new node is added, then storage capacity is expanded, but migration time increases and system performance decreases
Solution Approach 1:
The patent segments the data migration process by identifying and moving only the specific zones that need to be relocated to the new node, rather than migrating all data across the entire system. This selective zone-by-zone approach divides the migration task into manageable segments, reducing overall migration time while still achieving capacity expansion.
Solution Approach 2:
The patent performs preliminary actions by pre-calculating the stripe map before migration to identify exactly which zones need to be moved. This advance planning allows the system to prepare migration targets and minimize disruption during the actual migration process, reducing system performance impact.
2Quantity of substance
If data is re-striped across all nodes when a new node is added, then storage capacity is expanded, but system performance and availability decrease
Solution Approach 1:
The patent segments the migration workload to affect only specific zones rather than the entire data set. By isolating migration operations to individual zones that need relocation, other zones continue to serve I/O operations normally, maintaining system performance and availability during expansion.
Solution Approach 2:
The patent applies partial action by performing migration only on the necessary portions of data (specific zones) rather than all data. This selective approach achieves the required capacity expansion while minimizing the scope of disruptive operations, thereby preserving overall system performance.
3Quantity of substance
If massive data movement is performed during node addition, then new node capacity is utilized, but I/O operations are disrupted
Solution Approach 1:
The patent segments data into zones and migrates only those zones that need to be relocated to the new node. This selective zone migration minimizes the amount of data movement required, reducing I/O disruption while still achieving full utilization of the new node's capacity.
Solution Approach 2:
The patent maintains continuity of useful I/O operations by allowing non-migrated zones to continue serving requests during the migration process. This ensures that the system remains operational and reliable throughout the expansion, with only minimal disruption to affected zones.
Data Source
AI summary
A system, method, apparatus, and computer-readable medium are provided for expanding the data storage capacity of a virtualized storage system, such as a storage cluster. According to one method, maps are generated and stored that define a stripe pattern for storing data on the storage nodes of a storage cluster. The stripe pattern for each map is defined such that when a storage node is added to a cluster and the data is re-striped according to the new map, only the data that will subsequently reside in the new storage node is moved to the new storage cluster during re-striping. The stripe pattern may be further defined so that during re-striping no movement of data occurs between two storage nodes that existed in the cluster prior to the addition of the new storage node. The stripe pattern may be further defined such that during re-striping an equal amount of data is moved from each of the storage nodes that existed in the cluster prior to the addition of the new storage node to the new storage node.


