Affinity-Based Data Distribution in Mapped RAIN Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data storage techniques often result in underutilization of storage resources due to large storage groups, leading to inefficiencies in processor and network resource usage, especially when dealing with smaller data sets, as all disks in a node group are considered part of the same cluster, making it desirable to have more granular logical storage groups that can efficiently utilize larger real groups while providing smaller storage units for smaller data.

Innovation Solution

A mapped redundant array of independent nodes (RAIN) system that allows for more granular use of real clusters by defining logical storage locations across multiple hardware nodes, enabling data redundancy and flexibility in node addition or removal without data loss, using affinity metrics to distribute data across real nodes for optimal resource usage and accessibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If all disks of nodes in a group are considered part of the same cluster, then storage capacity is maximized, but storage granularity is reduced and resource efficiency deteriorates

Engineering Contradiction:
Improvestorage capacityVSAvoidstorage granularity
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent segments the storage system into multiple storage groups, each comprising a subset of disks from different nodes. This allows the large cluster to be divided into smaller logical units, providing both fine-grained storage allocation and efficient resource utilization. Each storage group can be independently managed and assigned to different workloads, resolving the contradiction between maximizing overall capacity and providing storage granularity.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If smaller groups with fewer nodes and disks are created, then storage granularity is improved, but processor and network resource efficiency deteriorates

Engineering Contradiction:
Improvestorage granularityVSAvoidresource efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent creates storage groups that can serve multiple purposes and workloads simultaneously. Each storage group is designed to be universally applicable for different data types and access patterns, allowing the system to maintain fine-grained storage allocation while efficiently utilizing processor and network resources across diverse operations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If data is distributed across multiple nodes, then data availability and redundancy are improved, but system complexity increases

Engineering Contradiction:
Improvedata availabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces storage group abstractions as intermediaries between the physical distributed storage infrastructure and the logical data access layer. This intermediary layer simplifies the complexity of managing data across multiple nodes by providing a unified interface for storage operations, while maintaining the reliability benefits of distributed storage through redundant data placement across storage groups.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11029865B2Affinity sensitive storage of data corresponding to a mapped redundant array of independent nodes
Publication Date: 2021.06.08 EMC IP HLDG CO LLC
  • US11029865B2 patent drawing
  • US11029865B2 patent drawing
  • US11029865B2 patent drawing

AI summary

Affinity sensitive storage of data corresponding to a mapped redundant array of independent nodes, e.g., mapped cluster, in a real storage system, e.g., a real cluster, is disclosed. Different mappings of mapped cluster data to real cluster storage locations can result in different levels of affinity between real nodes of the real cluster. A data storage scheme can be selected based on affinity scores, for example drawn from an affinity matrix, to provide access to stored data that can be more resilient against a real node becoming less available. Further, data recovery from a real node that has become less accessible can be improved where data is stored based on the affinity scores. Generally, data storage that provides greater diversity of data storage locations can be related to more desirable affinity scores. Further, data storage that provides less divergence of affinity scores across an affinity matrix can also be desirable.