Dispersed Storage Network Data Consistency via Dynamic Thresholds

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current dispersed storage networks face challenges in maintaining data consistency and integrity across geographically distributed storage units, particularly in scenarios where multiple storage units fail, leading to data loss or corruption.

Innovation Solution

The implementation of a dispersed storage network with a managing unit, integrity processing unit, and computing devices that utilize error encoding and decoding techniques, such as Cauchy Reed-Solomon encoding, to distribute data into encoded slices stored across multiple sites, ensuring data recovery and integrity through redundancy and secure storage protocols.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is distributed across multiple geographically dispersed storage units, then system availability and fault tolerance are improved, but data consistency and integrity become harder to maintain

Engineering Contradiction:
Improvefault toleranceVSAvoiddata consistency
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments data into multiple encoded slices distributed across geographically dispersed storage units. Each slice contains encoded information that contributes to the complete data set, allowing the system to tolerate failures of individual storage units while maintaining data consistency through the segmentation of information across multiple locations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback mechanisms where storage units report their status and data integrity to a coordinating system. This feedback loop enables the system to detect inconsistencies, identify failed storage units, and trigger reconstruction processes that restore data consistency by regenerating missing or corrupted slices from remaining valid slices.

Inventive Principle:
Principle #23Feedback

2Reliability

If error correction encoding is applied to distribute data across storage units, then data recovery capability is improved, but system complexity increases

Engineering Contradiction:
Improvedata recovery capabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies error correction encoding in advance during the data writing phase, creating redundant encoded slices before distribution. This preliminary action ensures that data recovery capability is built into the system architecture from the outset, allowing straightforward reconstruction of lost data without requiring complex real-time computation during failure scenarios.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates multiple copies of encoded data slices distributed across different storage units. These copies contain redundant information that enables data recovery, and the use of standardized copying mechanisms simplifies the overall system complexity compared to more sophisticated distributed consensus protocols.

Inventive Principle:
Principle #26Copying

3Reliability

If multiple storage units are used to store encoded data slices, then data security and redundancy are improved, but the difficulty of detecting and measuring data integrity increases

Engineering Contradiction:
Improvedata securityVSAvoiddata integrity verification
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent employs checksums and parity information that act as indicators of data integrity, similar to how color changes indicate state changes. Each encoded slice contains verification data that can be quickly checked to determine if the slice is intact or corrupted, making data integrity detection straightforward despite the distribution across multiple storage units.

Inventive Principle:
Principle #32Color changes

Data Source

PatentUS10402395B2Facilitating data consistency in a dispersed storage network
Publication Date: 2019.09.03 PURE STORAGE INC
  • US10402395B2 patent drawing
  • US10402395B2 patent drawing
  • US10402395B2 patent drawing

AI summary

A method for execution by a dispersed storage and task (DST) processing unit that includes a processor includes determining to access a set of storage units; identifying an information dispersal algorithm (IDA) width and a decode threshold number associated with the set of storage units; determining a number of available storage units of the set of storage units; determining a write threshold number and a read threshold number based on the number of available storage units and in accordance with a consistency approach; and accessing at least some of the available storage units utilizing at least one of the write and read threshold numbers.