Capacity-Leveling Data Distribution in Striped File Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In striped file systems, heterogeneous clusters with nodes of varying processor capabilities often lead to underutilization of processing resources, as newer nodes with additional computational power are not fully leveraged due to the need for homogeneous configurations, resulting in wasted processing power and increased costs.

Innovation Solution

A data distribution technique that uses a striping table to apportion data containers across volumes based on each node's capacity value, allowing for optimal assignment of stripes to nodes with varying capabilities, thereby maximizing the utilization of heterogeneous nodes within a striped volume set.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If nodes are configured to be homogeneous to ensure even data distribution, then data access performance is improved, but processing capacity and system scalability are limited

Engineering Contradiction:
Improvedata access performanceVSAvoidprocessing capacity
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by assigning different data distribution characteristics to different nodes based on their individual capacities. Each node receives a customized data distribution pattern that matches its processing capability, allowing heterogeneous nodes to function optimally without requiring uniform configuration across the cluster.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts data distribution parameters such as stripe sizes and distribution ratios based on node capacity values. By changing these parameters according to each node's processing capability, the system achieves both heterogeneous node utilization and optimal data access performance.

Inventive Principle:
Principle #35Parameter changes

2Power

If newer nodes with higher processor capabilities are added to the cluster, then processing power increases, but resource underutilization occurs due to homogeneous configuration requirements

Engineering Contradiction:
Improveprocessor capabilityVSAvoidresource utilization
Core Design Contradiction:
PowerVSProductivity

Solution Approach 1:

The patent enables local quality by tailoring data distribution patterns to match each node's specific capacity. Newer, more powerful nodes receive larger stripe assignments or different distribution ratios compared to older nodes, ensuring that each node operates at optimal utilization levels rather than being constrained by the weakest node in the cluster.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts the data distribution strategy when nodes are added or removed from the cluster. The data distribution manager recalculates optimal distribution patterns based on current node capacities, allowing the system to adapt to changing hardware configurations and maximize resource utilization.

Inventive Principle:
Principle #15Dynamics

3Device complexity

If data is striped evenly across all nodes, then data distribution is simple, but nodes with higher capacity are underutilized

Engineering Contradiction:
Improvedata distribution simplicityVSAvoidnode utilization
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent changes the data distribution parameters based on node capacity values. Instead of uniform distribution, the system calculates optimized distribution ratios that assign more data to higher-capacity nodes. This parameter adjustment maintains manageable complexity while significantly improving node utilization efficiency.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The data distribution manager automatically calculates and implements optimal data distribution patterns based on reported node capacities. The system self-adjusts without requiring manual intervention, using capacity information provided by nodes to autonomously determine optimal stripe assignments and distribution ratios.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8117388B2Data distribution through capacity leveling in a striped file system
Publication Date: 2012.02.14 NETAPP INC
  • US8117388B2 patent drawing
  • US8117388B2 patent drawing
  • US8117388B2 patent drawing

AI summary

A data distribution technique is configured to provide capacity leveling in a striped file system. When a new node is added to a striped volume set, the striping table is evolved to accommodate the newly added node. Each node of a cluster is illustratively associated with a capacity value that takes into account, e.g., processor speed, number of processors, hardware configuration and/or software available for the node. During the evolution process of the striping table, the technique apportions stripes of the SVS among the nodes in a manner so that they are optimally assigned to the nodes in accordance with each node's capacity value. By utilizing the evolutionary striping table that incorporates capacity values, heterogeneous nodes may be utilized to their maximum capacity within a striped volume set, thereby reducing underutilized processing resources.