Distributed Storage EC Redundancy Adjustment for Faulty Nodes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing distributed storage systems using erasure coding technology face performance impacts due to data migration and backward migration when a storage node becomes faulty, as they typically require adding a new node and migrating data, which affects system performance.
Innovation Solution
A data storage method and system that dynamically adjusts the EC redundancy ratio by allowing the storage client to store data blocks and parity blocks in a subset of nodes within a storage node group, excluding faulty nodes and reducing the number of generated EC blocks, thereby minimizing data migration and improving performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data migration and backward migration are performed to replace faulty storage nodes, then storage reliability is maintained, but system performance deteriorates due to extensive data redistribution
Solution Approach 1:
The patent dynamically changes the EC redundancy ratio parameter from a fixed value to a variable that adapts to the number of available storage nodes. When faulty nodes are detected, the system adjusts the redundancy ratio to match the reduced number of operational nodes, eliminating the need for data migration while maintaining storage reliability.
2Reliability
If the EC redundancy ratio is kept fixed, then storage reliability is ensured, but storage resource utilization decreases when nodes are faulty
Solution Approach 1:
The patent transforms the static EC redundancy ratio into a dynamic parameter that automatically adjusts based on the operational status of storage nodes. The system continuously monitors node health and modifies the redundancy ratio accordingly, enabling the storage system to adapt to changing conditions and optimize resource utilization while maintaining reliability.
Solution Approach 2:
The system changes the redundancy ratio parameter from a fixed configuration to a dynamically adjustable value. When storage nodes become faulty, the redundancy ratio is reduced to match the available capacity, preventing wasted storage resources while ensuring that reliability requirements are still met with the remaining nodes.
3Reliability
If new storage nodes are added to replace faulty nodes, then storage reliability is maintained, but device complexity increases due to node management overhead
Solution Approach 1:
Instead of adding physical nodes to maintain reliability, the system changes the redundancy ratio parameter to adapt to the reduced node count. This approach maintains reliability through parameter adjustment rather than hardware modification, significantly reducing the complexity of node management and system configuration.
Data Source
AI summary
A storage client needs to store to-be-written data into a distributed storage system, and storage nodes corresponding to a first data unit assigned for the to-be-written data by a management server are only some nodes in a storage node group. When receiving a status of the first data unit returned by the management server, the storage client may determine quantities of data blocks and parity blocks needing to be generated during EC coding on the to-be-written data. The storage client stores the generated data blocks and parity blocks into some storage nodes designated by the management server in a partition where the first data unit is located. Accordingly, dynamic adjustment of an EC redundancy ratio is implemented, and the management server may exclude some nodes in the partition from a storage range of the to-be-written data based on a requirement, thereby reducing a data storage IO amount.


