Grid Storage Server Responsibility Reassignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current fault-tolerant data storage systems, such as RAID arrays, face challenges in maintaining data integrity and availability due to the increasing likelihood of concurrent failures in larger disk arrays, and existing solutions do not adequately address the need for seamless redundancy and disaster recovery in virtualized grid storage systems.
Innovation Solution
A storage system comprising multiple disk units and a storage control grid with at least three data servers, where each server has direct or indirect access to the entire address space, with primary and secondary servers configured to handle requests and take over responsibilities in case of server shutdowns, enabling backward compatible hot upgrades and maintaining data protection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If RAID arrays use larger disk densities and more drives, then storage capacity increases, but the likelihood of concurrent failures increases
Solution Approach 1:
The storage system is segmented into multiple independent storage nodes, each capable of autonomous operation. Data is divided into chunks distributed across these nodes, with each node maintaining its own redundancy. This segmentation isolates failures to individual nodes, preventing concurrent failures from affecting the entire system, thus maintaining data protection while scaling capacity.
Solution Approach 2:
The patent introduces a new dimensional approach to redundancy by implementing hierarchical redundancy across multiple levels: within-node redundancy and cross-node redundancy. This multi-dimensional redundancy strategy allows the system to protect against concurrent failures by providing backup paths at different organizational levels, resolving the contradiction between increased capacity and maintained reliability.
2Reliability
If storage systems implement fault tolerance with multiple copies and parity, then data protection improves, but system complexity increases
Solution Approach 1:
Each storage node operates autonomously with self-managed redundancy and failure recovery capabilities. Nodes independently maintain their own data copies and parity information, and can autonomously recover from failures without requiring complex centralized coordination. This self-service approach simplifies the overall system architecture while maintaining robust data protection.
Solution Approach 2:
The storage nodes are designed as universal, multi-functional units that can perform multiple roles: primary storage, backup storage, and recovery source. Each node is capable of independently providing data protection services to multiple other nodes, eliminating the need for specialized redundant components and reducing overall system complexity while maintaining comprehensive data protection.
3Reliability
If storage systems use traditional RAID configurations, then data redundancy is achieved, but availability during server maintenance or failure is reduced
Solution Approach 1:
The system pre-distributes data copies and parity information across multiple storage nodes before any failure occurs. Each node maintains up-to-date redundant data in advance, so when a server needs maintenance or fails, the redundant copies are immediately available for takeover. This preliminary preparation ensures continuous availability without requiring complex real-time coordination during failure events.
Solution Approach 2:
The system implements dynamic role assignment where storage nodes can flexibly switch between primary and secondary roles based on operational needs. During maintenance or failure, nodes dynamically assume different responsibilities without requiring system reconfiguration. This dynamic adaptability maintains continuous availability while preserving data redundancy, resolving the contradiction between operational flexibility and data protection.
Data Source
AI summary
A method for hot backward compatible upgrade of a storage system includes: a) configuring each virtual partition (VP) to be controlled by a primary data server and a secondary data server b) configuring each data server to have primary responsibility over all logical block addresses (LBAs) corresponding to at least two virtual partitions and to have secondary responsibility over all LBAs corresponding to at least two other virtual partitions; c) responsive to a shut-down of a data server, i) re-configuring primary responsibility over each VP previously primary controlled by the shut-down server such that it becomes primary controlled by a server previously configured as a secondary server with respect to this VP; ii) re-allocating secondary responsibility over each VP previously secondary controlled by the shut-down server in a manner that each such VP becomes secondary controlled by a server other than the newly assigned server with primary responsibility.


