Tiered Storage Data Mover and Metadata Warehouse

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Lustre file systems face challenges in balancing storage capacity and IO throughput, leading to suboptimal performance and excessive costs, particularly in high-performance computing environments where existing storage solutions fail to meet performance requirements.

Innovation Solution

Implementing a system with front-end and back-end storage tiers, intermediate data mover modules, and a metadata warehouse to manage data movement and storage, enabling efficient data archiving, backup, and restoration while ensuring data integrity and compliance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If scale-out network attached storage is used for cost-effective storage, then storage capacity and cost are improved, but IO throughput and performance characteristics deteriorate

Engineering Contradiction:
Improvestorage capacityVSAvoidIO throughput
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The storage system is segmented into multiple storage tiers with different performance characteristics. High-performance storage tiers handle IO-intensive workloads while capacity-oriented tiers store less frequently accessed data, allowing the system to achieve both high capacity and high throughput by directing operations to appropriate tiers

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A data mover module acts as an intermediary between the storage tiers and the Lustre file system. This mediator intelligently moves data between tiers based on access patterns, caching frequently accessed data in high-performance tiers while maintaining capacity in cost-effective tiers, thus resolving the contradiction between capacity and throughput

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If storage devices are directly matched to current system needs, then IO performance is improved, but system flexibility and adaptability deteriorate

Engineering Contradiction:
ImproveIO performanceVSAvoidsystem flexibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The storage system employs dynamic data movement between tiers based on changing access patterns and workload requirements. The data mover continuously monitors and adjusts data placement, allowing the system to adapt to varying performance needs while maintaining a diverse portfolio of storage devices with different characteristics

Inventive Principle:
Principle #15Dynamics

3Quantity of substance

If data is moved between storage tiers, then storage capacity utilization is improved, but data movement costs and time increase

Engineering Contradiction:
Improvestorage capacity utilizationVSAvoiddata movement time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The system performs preliminary actions by proactively moving data between tiers based on predicted access patterns and policies. Frequently accessed data is pre-positioned in high-performance tiers before actual access occurs, reducing the time penalty of data movement while maintaining high capacity utilization in cost-effective tiers

Inventive Principle:
Principle #10Preliminary action

4Reliability

If metadata tracking is implemented for all data movement stages, then data integrity and validation are improved, but system complexity and overhead increase

Engineering Contradiction:
Improvedata integrityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The metadata warehouse implements self-service tracking where data movement information is automatically recorded and maintained without requiring complex external validation systems. The system serves its own metadata needs by having the data mover module autonomously update the warehouse with movement information, reducing overall system complexity while maintaining high data integrity

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9910742B1System comprising front-end and back-end storage tiers, data mover modules and associated metadata warehouse
Publication Date: 2018.03.06 EMC IP HLDG CO LLC
  • US9910742B1 patent drawing
  • US9910742B1 patent drawing
  • US9910742B1 patent drawing

AI summary

An information processing system comprises a plurality of front-end storage tiers, a plurality of back-end storage tiers, a plurality of data mover modules arranged between the front-end and back-end file storage tiers, and a metadata warehouse associated with the data mover modules and the front-end and back-end storage tiers. The data mover modules are configured to control movement of data between the storage tiers. The metadata warehouse is configured to store for each of a plurality of data items corresponding metadata comprising movement information characterizing movement of the data item between the storage tiers. The movement information for a given data item illustratively comprises locations, timestamps and checksums for different stages of movement of the given data item. Other types of metadata for the given data item illustratively include lineage information, access history information and compliance information.