Affinity-Based Data Distribution in Mapped RAIN Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage techniques often result in underutilization of storage resources due to large storage groups, leading to inefficiencies in processor and network resource usage, especially when dealing with smaller data sets, as all disks in a node group are considered part of the same cluster, making it desirable to have more granular logical storage groups that can efficiently utilize larger real groups while providing smaller storage units for smaller data.
Innovation Solution
A mapped redundant array of independent nodes (RAIN) system that allows for more granular use of real clusters by defining logical storage locations across multiple hardware nodes, enabling data redundancy and flexibility in node addition or removal without data loss, using affinity metrics to distribute data across real nodes for optimal resource usage and accessibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If all disks of nodes in a group are considered part of the same cluster, then storage capacity is maximized, but storage granularity is reduced and resource efficiency deteriorates
Solution Approach 1:
The patent segments the storage system into multiple storage groups, each comprising a subset of disks from different nodes. This allows the large cluster to be divided into smaller logical units, providing both fine-grained storage allocation and efficient resource utilization. Each storage group can be independently managed and assigned to different workloads, resolving the contradiction between maximizing overall capacity and providing storage granularity.
2Adaptability or versatility
If smaller groups with fewer nodes and disks are created, then storage granularity is improved, but processor and network resource efficiency deteriorates
Solution Approach 1:
The patent creates storage groups that can serve multiple purposes and workloads simultaneously. Each storage group is designed to be universally applicable for different data types and access patterns, allowing the system to maintain fine-grained storage allocation while efficiently utilizing processor and network resources across diverse operations.
3Reliability
If data is distributed across multiple nodes, then data availability and redundancy are improved, but system complexity increases
Solution Approach 1:
The patent introduces storage group abstractions as intermediaries between the physical distributed storage infrastructure and the logical data access layer. This intermediary layer simplifies the complexity of managing data across multiple nodes by providing a unified interface for storage operations, while maintaining the reliability benefits of distributed storage through redundant data placement across storage groups.
Data Source
AI summary
Affinity sensitive storage of data corresponding to a mapped redundant array of independent nodes, e.g., mapped cluster, in a real storage system, e.g., a real cluster, is disclosed. Different mappings of mapped cluster data to real cluster storage locations can result in different levels of affinity between real nodes of the real cluster. A data storage scheme can be selected based on affinity scores, for example drawn from an affinity matrix, to provide access to stored data that can be more resilient against a real node becoming less available. Further, data recovery from a real node that has become less accessible can be improved where data is stored based on the affinity scores. Generally, data storage that provides greater diversity of data storage locations can be related to more desirable affinity scores. Further, data storage that provides less divergence of affinity scores across an affinity matrix can also be desirable.


