Zero Information Gain Encoding for Secure Dispersed Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computer storage systems face challenges with data integrity and security due to the failure of memory devices, particularly those using physical movement technologies, such as disc drives, which can lead to data loss and increased maintenance demands, and RAID systems face inefficiencies and security risks with multiple copies of data.
Innovation Solution
A distributed storage system that uses error coding dispersal storage to encode data into multiple slices, which are then stored across geographically diverse locations, allowing for reliable and secure data retrieval even in the event of device failures, with a storage integrity processing unit verifying and rebuilding corrupted slices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is stored in multiple copies using RAID systems, then data reliability is improved, but storage capacity is reduced and security risks increase
Solution Approach 1:
The patent segments data into multiple slices using error coding (e.g., Reed-Solomon coding) where original data is divided into k slices and n-k parity slices are generated. This allows storing data across n locations without requiring full redundant copies, thus improving storage capacity utilization while maintaining data reliability through the ability to reconstruct data from any k slices.
Solution Approach 2:
The patent creates encoded copies of data slices that are distributed across multiple storage locations. Instead of storing identical full copies as in traditional RAID, the system stores encoded segments that can be combined to reconstruct the original data, reducing the total storage requirement while maintaining reliability.
2Reliability
If multiple copies of data are stored for redundancy, then data reliability is improved, but security risks worsen due to increased access points
Solution Approach 1:
By segmenting data into encoded slices distributed across multiple locations, the patent reduces security risks. Even if some slices are compromised, the segmented encoded structure makes it difficult to reconstruct the original data without the required threshold number of valid slices, thereby improving security while maintaining reliability.
3Quantity of substance
If disc drives with physical movement are used, then storage capacity is improved, but reliability worsens due to mechanical failures
Solution Approach 1:
The patent segments data across multiple storage devices, so that no single device failure results in data loss. The error coding scheme ensures that data can be reconstructed from remaining slices even if some devices fail, thus improving reliability while utilizing the storage capacity of mechanical drives.
4Reliability
If data is encoded and distributed across multiple locations, then data security and reliability are improved, but system complexity increases
Solution Approach 1:
The patent implements self-service through automatic encoding and decoding processes. The system automatically segments data into encoded slices, distributes them across storage locations, and reconstructs data when needed without requiring manual intervention, thus managing complexity through automation while improving reliability.
5Reliability
If RAID systems are used for redundancy, then data reliability is improved, but maintenance demands increase
Solution Approach 1:
The segmented encoded slice structure simplifies maintenance compared to traditional RAID. When a storage device fails, the system can reconstruct data by gathering the required number of slices from remaining devices and regenerating the missing slices, reducing maintenance complexity while maintaining reliability.
Data Source
AI summary
A method begins by a dispersed storage (DS) processing module encoding data using a dispersed storage error coding function to produce a set of encoded data slices. The method continues with the DS processing module encoding a first encoded data slice of the set of encoded data slices using a zero information gain (ZIG) function based on a second encoded data slice of the set of encoded data slices to produce a ZIG encoded data slice. The method continues with the DS processing module outputting the ZIG encoded data slice and a subset of encoded data slices of the set of encoded data slices, wherein the subset of encoded data slices includes less than a decode threshold number of encoded data slices and does not include the first or the second encoded data slice.


