Distributed Task Partitioning in Dispersed Storage Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current distributed storage networks face challenges in ensuring data integrity and security, particularly in the event of storage unit failures, where data loss can occur without redundant copies, and unauthorized access is a concern.

Innovation Solution

A dispersed storage network (DSN) utilizing error encoding techniques like Cauchy Reed-Solomon encoding disperses data across multiple storage units, allowing for data recovery even with failures and secure storage through encryption and secure encoding parameters managed by a managing unit, ensuring data integrity and security.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is stored in a distributed storage network without redundant copies, then storage efficiency is improved, but data reliability deteriorates when storage units fail

Engineering Contradiction:
Improvestorage efficiencyVSAvoiddata reliability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments data into multiple data slices and disperses them across different storage units in the network. This allows the system to store data efficiently without creating traditional redundant copies, while still maintaining reliability through the distributed nature of the slices.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies error encoding techniques that transform the data into an encoded form with specific parameters (such as Cauchy Reed-Solomon encoding). These parameter changes enable the system to recover original data even when some storage units fail, thus maintaining reliability without traditional redundancy.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If data is dispersed across multiple storage units using error encoding, then data reliability is improved, but system complexity increases

Engineering Contradiction:
Improvedata reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service mechanisms where the distributed storage network automatically performs error encoding, data slicing, and recovery operations without requiring external intervention. This manages the increased complexity through automation rather than manual processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system incorporates feedback mechanisms that monitor the state of storage units and automatically trigger appropriate actions (such as data recovery or re-encoding) when failures are detected. This feedback loop helps manage system complexity by responding to conditions in a structured manner.

Inventive Principle:
Principle #23Feedback

3Reliability

If traditional redundant copies are used for data storage, then data security is improved, but storage capacity utilization deteriorates

Engineering Contradiction:
Improvedata securityVSAvoidstorage capacity utilization
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Instead of creating traditional redundant copies that consume additional storage capacity, the patent segments data into slices and distributes them across the network. This segmentation approach provides security through distribution while maintaining efficient storage capacity utilization.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses error encoding parameter changes to create a system where data security is achieved through mathematical encoding rather than physical redundancy. This allows the system to maintain security while maximizing storage capacity utilization by storing only the essential data slices.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10303521B2Determining task distribution in a distributed computing system
Publication Date: 2019.05.28 PURE STORAGE INC
  • US10303521B2 patent drawing
  • US10303521B2 patent drawing
  • US10303521B2 patent drawing

AI summary

A method for execution by one or more processing modules of one or more computing devices of a dispersed storage network (DSN), by selecting a number of distributed storage and task execution (DST) EX units to favorably execute partial tasks of the corresponding tasks. The method continues by determining task partitioning based on one or more of distributed computing capabilities of the selected DST EX units. The method continues by determining processing parameters of the data based on the task partitioning. The method continues by partitioning the task(s) based on the task partitioning to produce the partial tasks. The method continues by processing the data in accordance with the processing parameters to produce slice groupings and sending the slice groupings and corresponding partial tasks to the DST EX units in accordance with a pillar mapping.