Multi-IDA Dispersed Storage for Lower I/O and Capacity Overhead

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional RAID systems face issues with effectiveness, efficiency, and security, particularly as the number of disks increases, leading to higher maintenance costs and risks of data loss due to disk failures, unauthorized access, and vulnerability to natural disasters.

Innovation Solution

A dispersed storage network (DSN) with error encoding and decoding capabilities, utilizing Cauchy Reed-Solomon encoding to distribute data across multiple storage units geographically, ensuring data integrity and security through redundancy and secure encoding parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple copies of data are stored in RAID systems to reduce data loss risk, then reliability is improved, but security deteriorates due to increased unauthorized access risk

Engineering Contradiction:
Improvedata loss preventionVSAvoidunauthorized access risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent segments data into multiple slices and distributes them across different storage units using information dispersal algorithms. Instead of storing complete copies of data, the system creates fragmented representations that require combination to reconstruct the original data, thereby improving security while maintaining reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces error encoding data as an intermediary layer between the original data and storage units. This encoding layer protects the underlying data structure and enables recovery from failures without exposing the original data, thus improving both reliability and security simultaneously.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If more disks are added to RAID array to increase storage capacity, then quantity of storage is improved, but reliability deteriorates due to higher disk failure probability

Engineering Contradiction:
Improvestorage capacityVSAvoiddisk failure risk
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent divides data into slices and distributes them across multiple storage units, allowing the system to scale storage capacity by adding more units without proportionally increasing failure risk. The segmentation allows selective access and recovery from subsets of storage units.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the fundamental parameter of data representation from complete copies to encoded slices. By using information dispersal algorithms with configurable redundancy levels, the system can adjust the relationship between storage capacity and reliability, allowing capacity to scale independently of failure probability.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If RAID systems store redundant copies of data, then reliability is improved, but device complexity increases due to maintenance requirements

Engineering Contradiction:
Improvedata protectionVSAvoidmaintenance complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-healing capabilities through error encoding and automatic data recovery mechanisms. When storage units fail, the system automatically detects the failure, retrieves data from remaining units using the encoding information, and reconstructs lost data without requiring manual intervention, thereby reducing maintenance complexity while maintaining reliability.

Inventive Principle:
Principle #25Self-service

4Reliability

If data is copied to multiple RAID devices for security, then reliability is improved, but loss of information increases due to potential unauthorized access

Engineering Contradiction:
Improvedata availabilityVSAvoiddata security
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments data into multiple slices that are distributed across storage units. Each slice alone is insufficient to reconstruct the original data, providing inherent security against unauthorized access while maintaining data availability through distributed storage and error encoding.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates encoded representations of data rather than direct copies. These encoded slices can be stored across multiple locations for availability, but they cannot be meaningfully interpreted without the complete encoded set, thus preventing information loss from unauthorized access while maintaining reliability.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10241694B2Reducing data stored when using multiple information dispersal algorithms
Publication Date: 2019.03.26 PURE STORAGE INC
  • US10241694B2 patent drawing
  • US10241694B2 patent drawing
  • US10241694B2 patent drawing

AI summary

Systems and methods for storing data in a dispersed storage network using at least two information dispersal algorithms (IDA' s) having different widths and thresholds are disclosed. In multiple IDA configurations, at least two IDA's with different widths and thresholds are paired and used to store the data multiple times, where some IDA's provide “wider” IDA configurations that are more reliable and other IDA's provide “narrower” configurations with a lower threshold and lower reliability. Data can be written in the less reliable IDA configurations as a performance optimization to reduce the input/output operations necessary for reading the data. As a further optimization, the processing unit can determine to write only a subset of the IDA configurations. Similarly, dispersed storage units themselves, when reaching the capacity limits for their memory devices, can begin to delete slices they hold for some of the IDA configurations, to free up space.