Dispersed Data Storage Grid Rebuild Mechanism
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data storage systems face challenges in securely and efficiently storing large amounts of data due to computational intensity of information dispersal algorithms, leading to vulnerabilities in unauthorized access and high costs associated with redundant storage solutions.
Innovation Solution
A distributed data file storage system using information dispersal algorithms to slice original data into subsets, which are then stored across multiple storage devices, allowing for secure and efficient reconstruction of data even with partial node outages, by employing coding algorithms and Rebuild Lists to manage data availability and integrity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If information dispersal algorithms are used to store data across multiple devices, then security and reliability are improved, but computational intensity increases
Solution Approach 1:
The patent divides original data into multiple dispersed data subsets using information dispersal algorithms, storing each subset on different storage devices. This segmentation approach enhances security and reliability while managing computational requirements through distributed processing across multiple nodes rather than centralized computation.
2Reliability
If dispersed data is stored across multiple storage devices, then unauthorized access is prevented, but data retrieval complexity increases
Solution Approach 1:
The patent implements a feedback mechanism where the system continuously monitors the operational status of storage devices and automatically manages data reconstruction. When devices are unavailable, the system detects this state and triggers automated reconstruction processes, reducing the perceived complexity for users while maintaining security.
Solution Approach 2:
The system performs automatic data reconstruction and redistribution without requiring manual intervention. When storage devices become unavailable, the system autonomously identifies missing data subsets, retrieves them from available nodes, and reconstructs the complete dataset, thereby simplifying the user experience despite the underlying complexity.
3Reliability
If redundant storage systems are used, then data reliability is improved, but storage costs increase
Solution Approach 1:
The patent transforms the storage model by changing the parameter of data representation through information dispersal algorithms. Instead of storing complete redundant copies of data, the system stores transformed dispersed subsets that require multiple nodes to reconstruct the original data, thereby reducing total storage requirements while maintaining reliability.
Solution Approach 2:
The system stores dispersed data subsets across multiple devices where each device holds only a partial portion of the complete dataset. This partial storage approach eliminates the need for full redundant copies on each device, reducing overall storage costs while ensuring data reliability through distributed storage and automatic reconstruction capabilities.
Data Source
AI summary
A digital data file storage system is disclosed in which original data files to be stored are dispersed using some form of information dispersal algorithm into a number of file “slices” or subsets in such a manner that the data in each file share is less usable or less recognizable or completely unusable or completely unrecognizable by itself except when combined with some or all of the other file shares. These file shares are stored on separate digital data storage devices as a way of increasing privacy and security. As dispersed file shares are being transferred to or stored on a grid of distributed storage locations, various grid resources may become non-operational or may operate below at a less than optimal level. When dispersed file shares are being written to a dispersed storage grid which not available, the grid clients designates the dispersed data shares that could not be written at that time on a Rebuild List. In addition when grid resources already storing dispersed data become non-available, a process within the dispersed storage grid designates the dispersed data shares that need to be recreated on the Rebuild List. At other points in time a separate process reads the set of Rebuild Lists used to create the corresponding dispersed data and stores that data on available grid resources.


