Storage Node Selection Using Encoded Data Preferences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current distributed storage systems face challenges in efficiently managing and retrieving large amounts of data across geographically dispersed locations, particularly in ensuring data integrity and availability while handling complex tasks, due to limitations in error correction and data partitioning strategies.

Innovation Solution

A distributed computing system that employs dispersed error encoding and decoding techniques, where data is segmented, encoded, and distributed across multiple execution units, allowing for reliable storage and retrieval of data and execution of tasks through pillar and slice groupings, ensuring data integrity and availability even with failures.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is distributed across multiple geographically dispersed locations, then data availability and reliability are improved, but data integrity and security become more difficult to ensure

Engineering Contradiction:
Improvedata availabilityVSAvoiddata integrity management
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments data into multiple slices and distributes them across different execution units in a dispersed storage network. Each slice is encoded with error correction information, allowing the system to maintain data integrity while distributing data across multiple locations. The segmentation principle is applied through slice grouping and pillar formation where data is divided into manageable units that can be independently stored and retrieved.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary encoding layer that processes data before distribution. Error correction codes and slice groupings act as intermediaries that protect data integrity during transmission and storage across dispersed locations. The encoding mechanism mediates between the original data and its distributed representation, ensuring that even if some slices are lost or corrupted, the original data can be reconstructed.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If error correction encoding is applied to distributed data, then data integrity is improved, but processing complexity and computational overhead increase

Engineering Contradiction:
Improvedata integrityVSAvoidencoding complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The error correction process is segmented into manageable stages: slicing data into smaller units, grouping slices into pillars, and applying encoding at the pillar level. This segmentation reduces the computational complexity of error correction by breaking down large data sets into smaller, independently processable units that can be encoded and decoded more efficiently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies error correction encoding selectively rather than to all data uniformly. Slice groupings allow the system to apply error correction only where needed, and the degree of error correction can be adjusted based on the importance and access patterns of different data slices. This partial application of error correction reduces overall processing complexity while maintaining data integrity for critical data.

Inventive Principle:
Principle #16Partial or excessive action

3Quantity of substance

If data is partitioned and distributed across multiple execution units, then storage scalability is improved, but data retrieval and task execution efficiency decrease

Engineering Contradiction:
Improvestorage capacityVSAvoiddata retrieval efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent performs preliminary actions by pre-grouping data slices into pillars and pre-encoding them with error correction information before distribution. This preliminary organization allows the system to retrieve data more efficiently because the grouping structure is already in place, eliminating the need for complex real-time reorganization during retrieval operations. Task execution units can directly access pre-grouped slices without additional processing overhead.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If redundant copies of data are stored, then data availability is improved, but storage space consumption increases

Engineering Contradiction:
Improvedata availabilityVSAvoidstorage space
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent uses error correction encoding as a form of intelligent copying rather than simple redundancy. Instead of storing multiple identical copies of the entire data set, the system creates encoded versions of data slices that can be combined to reconstruct the original data. This copying mechanism provides data availability similar to redundancy but consumes significantly less storage space because the encoded slices share information efficiently.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the parameter representation of data from raw storage to encoded storage. By transforming data into error correction encoded form, the system achieves better space efficiency while maintaining availability. The encoded representation allows the system to recover original data from fewer physical storage units compared to traditional redundancy approaches, effectively changing the storage density parameter.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11989093B2Selection of memory for data storage in a storage network
Publication Date: 2024.05.21 PURE STORAGE INC
  • US11989093B2 patent drawing
  • US11989093B2 patent drawing
  • US11989093B2 patent drawing

AI summary

Methods and apparatus for selection of memory devices in a distributed storage network. In an example, a computing device receives a data object for storage and selects a set of storage nodes of a plurality of sets of storage nodes for storing the data object. Selection of the set of storage nodes includes determining storage attributes associated with each set of storage nodes of the plurality of sets of storage nodes. Selection of the set of storage nodes additionally includes determining a storage preference associated with the data object, and comparing the storage preference with the storage attributes of the plurality of sets of storage nodes to determine a best match. Following selection of a set of storage nodes, the computing device facilitates storage of the data object in the selected set of storage nodes.