Selective Storage Volume Replication for Analytics Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data replication from business critical environments to analytics environments is time-consuming and impacts application performance, especially when sensitive information is involved, and existing remote replication methods do not adequately manage selective replication without compromising performance.

Innovation Solution

A storage system that replicates a subset of storage volumes based on user-specified database table IDs, using replication information to determine which regions can be replicated, allowing for selective and performance-optimized data transfer between storage systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If remote replication is executed on storage systems to copy data from business critical environment to analytics environment, then data availability is improved, but application performance deteriorates due to time-consuming replication process

Engineering Contradiction:
Improvedata availabilityVSAvoidapplication performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The storage volume is divided into multiple regions, and replication is performed selectively for specific regions rather than the entire volume. This segmentation allows only necessary data to be replicated, reducing the replication overhead and its impact on application performance while maintaining data availability for analytics workloads.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different replication policies are applied to different regions of the storage volume based on their specific requirements. Business critical regions may have different replication settings compared to analytics regions, allowing optimized performance for each area while achieving overall data availability goals.

Inventive Principle:
Principle #3Local quality

2Productivity

If selective replication of storage regions is implemented to minimize performance impact, then application performance is improved, but device complexity increases due to replication information management

Engineering Contradiction:
Improveapplication performanceVSAvoidreplication information management
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The storage system automatically manages replication information without requiring manual intervention. The system self-configures which regions to replicate based on predefined policies and automatically updates replication status, reducing the operational complexity despite the selective replication requirements.

Inventive Principle:
Principle #25Self-service

3Reliability

If complete volume replication is performed to ensure data availability, then reliability is improved, but time consumption increases making the process inefficient

Engineering Contradiction:
Improvedata availabilityVSAvoidreplication time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

Only the necessary portions of data (specific regions corresponding to analytics workloads) are extracted and replicated from the source volume, rather than replicating the entire volume. This extraction approach significantly reduces replication time while ensuring data availability for the analytics environment.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11196806B2Method and apparatus for replicating data between storage systems
Publication Date: 2021.12.07 HITACHI VANTARA LTD
  • US11196806B2 patent drawing
  • US11196806B2 patent drawing
  • US11196806B2 patent drawing

AI summary

Example implementations described herein are directed to replication of data between different environments selectively while maintaining the performance of applications. The replication of data may be used to replicate data from a data center running a business critical application to another data center running an analytics application. In example implementations, a storage management program translates the IDs of database tables to storage locations, and the storage management program requests a storage system to replicate those storage locations to another storage system.