Hardware Controller Parity Storage Container Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing parity-based storage systems face challenges in efficiently recovering data when storage containers fail, particularly in multi-node environments, due to redundant computation and inefficient inter-node communication.
Innovation Solution
The system employs a hardware controller to identify failed storage containers, recover data using parity containers from other nodes, and store the recovered data in spare containers, optimizing inter-node communication by avoiding redundant calculations and data copying.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional parity-based recovery methods are used in multi-node storage systems, then data can be recovered from failed storage containers, but redundant computation and inefficient inter-node communication occur
Solution Approach 1:
The patent segments the recovery process by identifying which storage containers have failed and selectively recovering only those specific containers rather than performing blanket recovery operations across all nodes. The hardware controller divides the multi-node system into targeted recovery units based on failure locations, enabling parallel processing of independent recovery tasks and reducing overall computation time.
Solution Approach 2:
The system performs preliminary identification of failed storage containers before initiating recovery operations. The hardware controller scans and detects failed containers upfront, then uses this information to optimize the recovery sequence and data retrieval paths, avoiding unnecessary computation on already-recovered or non-failed containers.
2Reliability
If comprehensive data recovery is performed across all storage nodes, then data integrity is maintained, but network communication overhead and computation time increase
Solution Approach 1:
The patent applies local quality by tailoring the recovery operation to the specific failure scenario. Instead of uniform recovery across all nodes, the hardware controller applies different recovery strategies based on which specific containers failed and where they are located. This localized approach maintains data integrity for affected containers while avoiding unnecessary operations on healthy data, reducing overall recovery time.
3Reliability
If parity containers are accumulated and copied across multiple nodes, then data can be reconstructed, but redundant data copying and computation occur
Solution Approach 1:
The hardware controller performs preliminary identification of failed storage containers before initiating recovery operations. By detecting failures upfront and using this information to optimize the recovery sequence, the system avoids unnecessary computation on already-recovered or non-failed containers, reducing overall complexity.
Solution Approach 2:
The recovery process is segmented into targeted operations based on failure locations. The system divides the multi-node system into independent recovery units, allowing parallel processing of discrete recovery tasks and reducing the overall computational complexity compared to blanket recovery approaches.
Data Source
AI summary
Systems and methods for recovery of parity based storage systems are described. In one embodiment, a group of nodes includes one or more storage nodes, and the one or more storage nodes include one or more storage containers. In one embodiment, the one or more storage containers include one or more data storage containers, one or more parity storage containers, or one or more spare storage containers, or any combination thereof. The system and methods include a hardware controller configured to identify a first failed storage container on a first storage node from the group of storage nodes, identify data associated with the first failed storage container on at least a second storage container on a second storage node from the plurality of storage nodes, and recover the data associated with the first failed storage container from at least the second storage container on the second storage node.


