Dispersed Storage Network Data Migration for Cohort Scaling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current dispersed storage networks face challenges in efficiently managing and maintaining data integrity across multiple storage units, particularly in scenarios where storage units fail or data corruption occurs, leading to potential data loss without redundant copies.

Innovation Solution

The implementation of a dispersed storage network (DSN) with a managing unit, integrity processing unit, and computing devices that utilize error encoding and decoding techniques, such as Cauchy Reed-Solomon encoding, to distribute data across multiple storage units, ensuring data recovery and security through encoded data slices and decentralized agreement modules for resource allocation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is distributed across multiple storage units without redundant copies, then storage efficiency is improved, but data reliability deteriorates when storage units fail

Engineering Contradiction:
Improvestorage efficiencyVSAvoiddata reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments data into multiple encoded slices distributed across different storage units. Each slice contains encoded information that contributes to the whole, allowing data to be reconstructed from any sufficient subset of slices. This segmentation enables storage efficiency by eliminating redundant copies while maintaining reliability through error encoding.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of data representation by applying error encoding transforms to the original data before distribution. This transformation creates encoded slices with specific mathematical properties that enable reconstruction. The parameter change from raw data to encoded slices allows the system to achieve both storage efficiency and data reliability simultaneously.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If error encoding is applied to distribute data across storage units, then data security is improved, but processing complexity increases

Engineering Contradiction:
Improvedata securityVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates multiple encoded copies of data segments through error encoding, where each copy (slice) contains redundant information in encoded form. These encoded copies enable data recovery without requiring identical redundant copies, thus improving data security while managing processing complexity through efficient encoding algorithms.

Inventive Principle:
Principle #26Copying

3Quantity of substance

If storage capacity is increased by adding more storage units, then data availability is improved, but system complexity increases

Engineering Contradiction:
Improvestorage capacityVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent creates a universal encoded data structure that can be distributed across any number of storage units without changing the fundamental encoding scheme. The same error encoding process works whether data is distributed across 3 or 30 storage units, providing multi-functionality that increases storage capacity while minimizing system complexity through a unified approach.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10387070B2Migrating data in response to adding incremental storage resources in a dispersed storage network
Publication Date: 2019.08.20 PURE STORAGE INC
  • US10387070B2 patent drawing
  • US10387070B2 patent drawing
  • US10387070B2 patent drawing

AI summary

A method for execution by a computing device includes detecting that an incremental storage cohort has been added to a storage generation to produce an updated plurality of storage cohorts of an updated storage generation, where each storage cohort includes a set of storage units. For each storage cohort, a slice listing process is initiated to identify a plurality of DSN addresses associated with storage of data objects within the each storage cohort. For each DSN address, ranked scoring information is obtained for the each storage cohort of the updated plurality of storage cohorts. One storage cohort is identified based on the ranked scoring information. When the identified storage cohort is different than another storage cohort associated with current storage of encoded data slices associated with the DSN address of the identified storage cohort, a migration process is initiated to migrate the encoded data slices to the identified storage cohort.