NVMe Stream ID Generation via Replication Statistics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data storage replication systems face challenges in optimizing data placement and leveraging Non-Volatile Memory Express (NVMe) stream capabilities due to reliance on stream information from the replication source and additional metadata storage requirements.
Innovation Solution
The solution involves automatically generating NVMe stream IDs based on replication IO statistics, allowing for optimal data placement without relying on source stream information and eliminating the need for additional metadata storage, thereby enabling efficient data storage optimization across replication target objects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If stream information from replication source is used for data placement, then data placement optimization is achieved, but additional metadata storage is required
Solution Approach 1:
The patent extracts the stream identification functionality from the source system and implements it autonomously at the target system using replication statistics. This removes the dependency on source stream information and eliminates the need to store additional metadata about stream IDs, while still achieving data placement optimization through locally-generated stream identifiers
Solution Approach 2:
The target system serves itself by automatically generating NVMe stream identifiers based on its own observation of replication statistics. Instead of relying on the source system to provide stream information, the target system independently creates stream IDs by hashing replication metrics, thereby eliminating the need for additional metadata storage while maintaining data placement capabilities
2Productivity
If NVMe stream capabilities are fully leveraged, then data storage optimization is improved, but complexity in stream ID management increases
Solution Approach 1:
The patent transforms the stream ID generation process by changing the input parameters from source-provided stream identifiers to target-generated identifiers based on replication statistics. This parameter change simplifies management by using readily available replication metrics (which already exist for performance monitoring) rather than requiring separate stream ID tracking mechanisms
Solution Approach 2:
The patent makes replication statistics serve multiple functions: they continue to provide performance monitoring information while simultaneously serving as the basis for generating NVMe stream identifiers. This multi-functionality eliminates the need for separate stream ID management systems, reducing complexity while still enabling full leverage of NVMe stream capabilities for data storage optimization
Data Source
AI summary
An aspect of optimizing storage of data in a data replication system includes, for a plurality of write requests received from a source site, determining transfer statistics corresponding to each of the write requests and updating a table with the transfer statistics. An aspect also includes grouping pages in the table having common transfer statistics, assigning a unique non-volatile memory express (NVMe) stream identifier (ID) to each of the groups, and identifying grouped pages based on the assigned NVMe stream ID. An aspect further includes selecting a storage optimization technique for each of the groups based on the common transfer statistics and storing data of the write requests for each of the groups according to the selected optimization technique.


