Dispersed Storage Network Data Slicing for Reliability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage systems, such as RAID, face challenges with disk drive failures leading to data loss, especially as the number of disks increases, and they are vulnerable to unauthorized access when mirroring data across locations.
Innovation Solution
A dispersed storage network (DSN) system that breaks data into error-coded slices and disperses them across multiple storage units, using error encoding and sub-slicing to ensure data integrity and security, with a management system that optimizes storage and retrieval across diverse storage units.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If RAID uses more disk drives to increase storage capacity, then storage capacity is improved, but the probability of data loss increases due to higher failure rates
Solution Approach 1:
The patent segments data into multiple slices and disperses them across many storage units rather than using traditional RAID blocks. This segmentation allows the system to tolerate individual unit failures while maintaining data integrity, as long as the minimum threshold of slices is recovered.
Solution Approach 2:
The patent introduces error correction codes as an intermediary layer between the data and storage units. These codes enable the system to recover from failures without requiring exact replicas of the original data, thus improving reliability while maintaining storage efficiency.
2Reliability
If RAID mirrors data across multiple locations to reduce data loss risk, then reliability is improved, but vulnerability to unauthorized access increases
Solution Approach 1:
By segmenting data into slices and dispersing them across multiple locations with different operating systems, the patent prevents unauthorized access. Even if one location is compromised, the attacker cannot reconstruct the full data without obtaining the minimum threshold of slices from diverse locations.
Solution Approach 2:
The patent assigns different operating systems to different storage units in the array. This local quality diversity creates heterogeneous environments that are more resistant to unauthorized access, as a single exploitation method cannot compromise all units simultaneously.
3Ease of manufacture
If RAID uses lower quality disk drives to reduce cost, then device cost is reduced, but the probability of failure increases
Solution Approach 1:
The patent accepts that individual storage units may fail and design the system accordingly. Rather than using expensive high-reliability drives, the system uses cheaper units with the understanding that some will fail, but error correction codes ensure data recovery as long as the minimum threshold is maintained.
Solution Approach 2:
Error correction codes serve as an intermediary that compensates for the lower reliability of individual units. These codes enable the system to tolerate failures of cheaper drives while maintaining overall data integrity.
4Reliability
If conventional storage systems use redundant disk drives to protect against failure, then reliability is improved, but storage efficiency decreases due to overhead
Solution Approach 1:
The patent changes the redundancy parameter from traditional RAID parity overhead to dispersed storage with error correction codes. This allows the system to achieve comparable or superior reliability with lower overhead by distributing data across more units rather than using duplicate copies.
Solution Approach 2:
By segmenting data into slices and distributing them across many units, the patent reduces the overhead required for redundancy. Instead of dedicating entire drives to parity or mirror copies, the system uses a fraction of each unit's capacity for error correction, improving overall storage efficiency.
Data Source
AI summary
A plurality of data slices are generated from a block of data to be stored in the dispersed storage system. A plurality of dispersed storage units are determined for storing the plurality of data slices, based on a geographical location associated with the plurality of dispersed storage units.


