Heterogeneous Memory Erasure Coding Rebuild
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Solid-state drives designed to conform to hard disk drive standards often fail to leverage the unique characteristics of flash and other solid-state memories, limiting their ability to provide enhanced features and data redundancy.
Innovation Solution
A method for proactively rebuilding user data across multiple storage nodes within a single chassis using erasure coding, allowing the system to remain operational even if two nodes fail, by distributing data and metadata and employing differing erasure coding schemes for reading and writing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If solid-state drives conform to hard disk drive standards for compatibility, then compatibility is improved, but the ability to leverage unique flash memory characteristics and provide enhanced data redundancy is worsened
Solution Approach 1:
The system dynamically selects between different erasure coding schemes (RS, BCH, LDPC) based on the specific storage scenario and requirements. This dynamic adaptation allows the system to optimize for either compatibility or enhanced redundancy depending on the operational context, resolving the contradiction between conforming to standards and leveraging unique flash characteristics
Solution Approach 2:
The patent changes the parameter of erasure coding scheme selection based on storage conditions. By having multiple codable redundancy schemes available and selecting appropriate ones, the system can adjust its redundancy approach to match the specific characteristics of flash memory while maintaining compatibility where needed
2Reliability
If data is distributed across multiple storage nodes with erasure coding, then data redundancy and availability are improved, but system complexity increases
Solution Approach 1:
The system segments data into multiple shards distributed across different storage nodes, with redundancy information also segmented and distributed. This segmentation approach provides fault tolerance while keeping individual node complexity manageable, as each node only needs to handle its specific data segments rather than the entire system
Solution Approach 2:
The patent introduces a coordination mechanism that acts as an intermediary to manage the complexity of multi-node erasure coding operations. This intermediary handles the coordination of read/write operations across nodes, managing the complexity centrally while allowing individual nodes to remain relatively simple
3Reliability
If proactive data rebuilding is performed when storage nodes become unreachable, then data integrity is improved, but processing time and system overhead increase
Solution Approach 1:
The system performs preliminary actions by detecting unreachable storage nodes and initiating data rebuilding proactively before data requests fail. This preliminary detection and rebuilding approach prevents data loss and avoids the need for reactive recovery operations that would take longer and disrupt service
Solution Approach 2:
The patent implements a feedback mechanism where the system continuously monitors storage node availability and uses this feedback to trigger rebuilding operations. When nodes become unreachable, the feedback loop initiates proactive rebuilding, optimizing the balance between data integrity and processing time by acting at the optimal moment
Data Source
AI summary
A method for proactively rebuilding user data in a plurality of storage nodes of a storage cluster is provided. The method includes distributing user data and metadata throughout the plurality of storage nodes such that the plurality of storage nodes can read the user data, using erasure coding, despite loss of two of the storage nodes. The method includes determining that one of the storage nodes is unreachable and determining to rebuild the user data for the one of the storage nodes that is unreachable. The method includes reading the user data across a remainder of the plurality of storage nodes, using the erasure coding and writing the user data across the remainder of the plurality of storage nodes, using the erasure coding. A plurality of storage nodes within a single chassis that can proactively rebuild the user data stored within the storage nodes is also provided.


