NVMe Stream ID Generation via Replication Statistics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data storage replication systems face challenges in optimizing data placement and leveraging Non-Volatile Memory Express (NVMe) stream capabilities due to reliance on stream information from the replication source and additional metadata storage requirements.

Innovation Solution

The solution involves automatically generating NVMe stream IDs based on replication IO statistics, allowing for optimal data placement without relying on source stream information and eliminating the need for additional metadata storage, thereby enabling efficient data storage optimization across replication target objects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If stream information from replication source is used for data placement, then data placement optimization is achieved, but additional metadata storage is required

Engineering Contradiction:
Improvedata placement optimizationVSAvoidmetadata storage
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent extracts the stream identification functionality from the source system and implements it autonomously at the target system using replication statistics. This removes the dependency on source stream information and eliminates the need to store additional metadata about stream IDs, while still achieving data placement optimization through locally-generated stream identifiers

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The target system serves itself by automatically generating NVMe stream identifiers based on its own observation of replication statistics. Instead of relying on the source system to provide stream information, the target system independently creates stream IDs by hashing replication metrics, thereby eliminating the need for additional metadata storage while maintaining data placement capabilities

Inventive Principle:
Principle #25Self-service

2Productivity

If NVMe stream capabilities are fully leveraged, then data storage optimization is improved, but complexity in stream ID management increases

Engineering Contradiction:
Improvedata storage optimizationVSAvoidstream ID management
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent transforms the stream ID generation process by changing the input parameters from source-provided stream identifiers to target-generated identifiers based on replication statistics. This parameter change simplifies management by using readily available replication metrics (which already exist for performance monitoring) rather than requiring separate stream ID tracking mechanisms

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent makes replication statistics serve multiple functions: they continue to provide performance monitoring information while simultaneously serving as the basis for generating NVMe stream identifiers. This multi-functionality eliminates the need for separate stream ID management systems, reducing complexity while still enabling full leverage of NVMe stream capabilities for data storage optimization

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10664268B2Data storage optimization using replication statistics to automatically generate NVMe stream identifiers
Publication Date: 2020.05.26 EMC IP HLDG CO LLC
  • US10664268B2 patent drawing
  • US10664268B2 patent drawing
  • US10664268B2 patent drawing

AI summary

An aspect of optimizing storage of data in a data replication system includes, for a plurality of write requests received from a source site, determining transfer statistics corresponding to each of the write requests and updating a table with the transfer statistics. An aspect also includes grouping pages in the table having common transfer statistics, assigning a unique non-volatile memory express (NVMe) stream identifier (ID) to each of the groups, and identifying grouped pages based on the assigned NVMe stream ID. An aspect further includes selecting a storage optimization technique for each of the groups based on the common transfer statistics and storing data of the write requests for each of the groups according to the selected optimization technique.