Dual Network Storage Cluster Architecture for Node Failure Resilience
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current storage systems face challenges in efficiently communicating and managing data across multiple storage nodes and units, particularly in maintaining data availability and redundancy without relying on a single node, and in dynamically reconfiguring storage capacity and workload distribution.
Innovation Solution
A storage cluster architecture with a first network connecting storage nodes and a second network connecting storage units, enabling proactive data rebuilding and redundancy management, where storage nodes and units can communicate directly to maintain data integrity and availability, and allowing for flexible reconfiguration and expansion of storage capacity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single storage node is used to manage all storage units, then the system structure is simple, but the system reliability deteriorates when the storage node fails
Solution Approach 1:
The patent segments the centralized storage management function into distributed peer-to-peer interactions among storage units. Each storage unit can directly communicate with others on the fabric, eliminating the single point of failure while maintaining manageable complexity through standardized communication protocols
Solution Approach 2:
The fabric acts as an intermediary communication medium that enables direct storage unit-to-storage unit communication without requiring a centralized storage node, thereby improving reliability while keeping the overall system structure relatively simple
2Device complexity
If storage units communicate through a centralized storage node, then the communication protocol is simple, but the communication efficiency deteriorates
Solution Approach 1:
The communication path is segmented into direct peer-to-peer connections between storage units on the fabric, eliminating the intermediate centralized node and reducing communication hops, thereby improving efficiency while maintaining protocol simplicity through standardized fabric protocols
Solution Approach 2:
The patent introduces a new communication dimension by enabling direct lateral communication between storage units on the fabric, rather than forcing all communication through the vertical centralized node path, thus improving efficiency without significantly increasing protocol complexity
3Stability of the object's composition
If fixed storage capacity is allocated to each node, then the system is stable, but the adaptability deteriorates when reconfiguration is needed
Solution Approach 1:
The patent implements dynamic storage capacity allocation where storage units can flexibly request and receive additional capacity from the fabric pool as needed, allowing the system to adapt to changing requirements while maintaining stable operation through formal capacity request and allocation protocols
Data Source
AI summary
A storage system is provided. The storage system includes a plurality of storage nodes, each of the plurality of storage nodes having a plurality of storage units with storage memory. The system includes a first network coupling the plurality of storage nodes and a second network coupled to at least a subset of the plurality of storage units of each of the plurality of storage nodes such that one of the plurality of storage units of a first one of the plurality of storage nodes can initiate or relay a command to one of the plurality of storage units of a second one of the plurality of storage nodes via the second network without the command passing through the first network.


