Dynamic Resource Balancing in Asynchronous Replication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data storage systems face inefficiencies in achieving low Recovery Point Objective (RPO) for asynchronous replication, particularly due to resource-intensive snapshot difference techniques and high overhead costs.
Innovation Solution
The system configures stretched volumes for asynchronous replication with dynamic service level adjustments based on resource consumption, utilizing write tracking, transient snapshots, and caching to optimize data replication, thereby achieving near-zero or low RPO.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If snapshot difference techniques are used for asynchronous replication, then data consistency is improved, but resource consumption increases
Solution Approach 1:
The patent segments the replication process into multiple service levels (e.g., high, medium, low priority) that can be dynamically adjusted. Different stretched volumes are assigned to different service levels based on their RPO requirements, allowing resource-intensive snapshot operations to be concentrated on critical volumes while using lighter replication methods for less critical ones.
Solution Approach 2:
The system dynamically adjusts the replication service level for each stretched volume based on current resource availability and RPO compliance status. When resources are abundant, higher service levels with more frequent snapshots are used. When resources are constrained, the system downgrades service levels to reduce resource consumption while maintaining RPO targets.
2Manufacturing precision
If frequent snapshots are taken for low RPO, then replication accuracy is improved, but overhead costs increase
Solution Approach 1:
Different snapshot frequencies and service levels are applied locally to different stretched volumes based on their individual RPO requirements and criticality. Instead of uniformly applying high-frequency snapshots to all volumes, the system tailors the replication intensity to each volume's specific needs, reducing overall overhead while maintaining required accuracy for critical data.
Solution Approach 2:
The system changes operational parameters (snapshot interval, service level priority) dynamically based on resource conditions and RPO compliance. When RPO targets are met with lower resource usage, the system reduces snapshot frequency. When compliance is at risk or resources are abundant, parameters are adjusted to increase replication accuracy.
3Reliability
If multiple stretched volumes are replicated simultaneously, then data availability is improved, but resource contention increases
Solution Approach 1:
The system dynamically prioritizes replication operations across multiple stretched volumes using service levels. When resource contention is detected, lower-priority volumes temporarily reduce their snapshot frequency or defer replication operations, allowing critical high-priority volumes to maintain their replication schedule and RPO compliance.
Solution Approach 2:
The system monitors resource consumption and RPO compliance status for each stretched volume and uses this feedback to adjust replication priorities and service levels. When resource contention prevents RPO compliance for certain volumes, the system identifies which volumes can temporarily accept higher RPO values and adjusts their service levels accordingly to relieve pressure on critical systems.
Data Source
AI summary
Techniques can include: configuring stretched volumes for asynchronous replication; specifying replication settings denoting selected replication service levels of asynchronous replication optimizations for the stretched volumes; performing asynchronous replication for the stretched volumes based on the replication settings; monitoring a current amount denoting an amount of a resource that is free and available for use; and responsive to determining that the current amount is below a minimum, performing a corrective action to increase the current amount, wherein the corrective action includes: changing a replication setting for a stretched volume from a first replication service level to a second replication service level, wherein the second replication service level is expected to consume less of the resource than the first replication service level when performing asynchronous replication for the stretched volume.


