Resilient Storage Nodes with Equal Cell Distribution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional direct attached storage (DAS) systems face limitations in data accessibility and failure recovery, as data access is lost when a server fails, whereas storage area networks (SANs) provide redundancy but are complex to manage effectively.
Innovation Solution
A storage system with equal non-volatile storage capacity across nodes, subdivided into equal cells, and protection groups distributed across nodes such that no more than one member of any group is stored on a single node, allowing for efficient rebuilding of data on remaining nodes in case of failure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional DAS systems are used, then data access is simple and direct, but data accessibility is lost when a server fails
Solution Approach 1:
The storage system is segmented into multiple independent storage nodes, each capable of storing members of protection groups. This segmentation allows the system to maintain data accessibility even when individual nodes fail, as data can be retrieved from remaining nodes. The storage capacity of each node is further subdivided into equal-sized cells for organized data placement.
Solution Approach 2:
Each storage node is assigned specific members of protection groups based on deterministic algorithms, creating local quality differences across nodes. Each node stores a unique subset of data members, ensuring that no single node contains all members of any protection group. This local differentiation enables fault tolerance while maintaining overall system reliability.
2Reliability
If SAN architecture is used to provide redundancy, then data accessibility is maintained during failures, but management complexity increases
Solution Approach 1:
The system changes the organizational parameters of data storage by using protection groups with distributed members across nodes. Each protection group has members stored in cells across multiple nodes, with deterministic placement rules. This parameter-based organization simplifies management compared to traditional SAN architectures, as the deterministic algorithms automatically handle data placement and recovery without complex manual configuration.
Solution Approach 2:
The storage system implements self-service capabilities through deterministic algorithms that automatically manage data placement, distribution, and recovery. When a storage node fails, the system automatically rebuilds data on remaining nodes using the protection group structure, without requiring manual intervention. This self-managing approach reduces operational complexity while maintaining high availability.
3Reliability
If data is distributed across multiple nodes, then resiliency is improved, but data movement increases during rebuilding
Solution Approach 1:
The system performs preliminary action by pre-distributing members of protection groups across multiple storage nodes using deterministic algorithms before any failure occurs. Each node is pre-configured with specific cells for storing protection group members, establishing a ready-made recovery structure. This preliminary distribution minimizes data movement during failures, as the system only needs to relocate data to pre-designated cells on remaining nodes rather than performing complex real-time redistribution.
4Productivity
If storage capacity is equalized across nodes, then resource utilization is optimized, but flexibility in data placement is reduced
Solution Approach 1:
The storage system achieves universality by designing all storage nodes with equal capacity and identical cell structures, allowing any node to serve any function within the protection group structure. Each node can store members from any protection group, and the deterministic algorithms can dynamically assign members to any available node. This universal design maximizes space efficiency while maintaining placement flexibility, as the same node structure can adapt to different data distribution patterns based on system needs.
Data Source
AI summary
A storage system has a plurality of storage nodes having equal non-volatile storage capacity that is subdivided into equal size cells. Host application data that is stored in the cells is protected using RAID or EC protection groups each having members stored in ones of the cells and distributed across the storage nodes such that no more than one member of any single protection group is stored by any one of the storage nodes. Spare cells are maintained for rebuilding protection group members of a failed one of the storage nodes on remaining non-failed storage nodes so full data access is possible before replacement or repair of the failed storage node.


