Dynamic Secondary Storage Resource Scaling for Failover RTO Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing remote copying systems face challenges in preventing recovery requirement violations, such as RTO, during failovers to secondary sites while maintaining cost-effectiveness in hardware configurations for secondary sites in normal times.
Innovation Solution
A computer system with a primary and secondary site storage system connected via a network, where a management device can dynamically change resources at the secondary site, performing failover and resource enhancement to ensure compliance with recovery requirements without excessive hardware provisioning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If sufficient hardware is prepared at the secondary site to withstand both remote copying process and host I/O process after failover, then recovery requirement (RTO) is satisfied, but hardware cost in normal times becomes excessive
Solution Approach 1:
The patent implements dynamic resource allocation where the secondary site hardware configuration changes based on operational state. In normal times, only minimal hardware sufficient for remote copying is provisioned. Upon failover detection, the system dynamically scales up resources to accommodate full host I/O workloads, and scales down after failback, ensuring RTO compliance without permanent excessive hardware provisioning
Solution Approach 2:
The system performs preliminary resource enhancement before actual failover occurs. When a failure is detected at the primary site, the management device proactively enhances resources at the secondary site before the failover is executed, ensuring that sufficient capacity is already available when the takeover begins, thus meeting RTO requirements
2Quantity of substance
If hardware configuration at secondary site is minimized to withstand only remote copying process, then hardware cost is reduced, but recovery requirement may be violated during failover
Solution Approach 1:
The system transitions from a static minimal hardware configuration to a dynamic configuration that adapts to operational demands. The management device monitors system state and automatically adjusts resource allocation at the secondary site, scaling up capacity when failover is required and scaling down when not needed, thus maintaining cost-effectiveness while ensuring recovery requirements are met when necessary
Solution Approach 2:
The management device performs preliminary resource enhancement before failover execution. When a primary site failure is detected, the system proactively increases hardware resources at the secondary site before the actual failover occurs, ensuring sufficient capacity is available to handle the incoming workload and meet RTO commitments
3Quantity of substance
If dynamic hardware change is implemented at secondary site to match operational needs, then hardware cost is optimized, but change time may cause RTO violation
Solution Approach 1:
The system executes resource enhancement in advance before the actual failover operation. When a primary site failure is detected, the management device immediately triggers resource scaling at the secondary site before the failover switch occurs. This preliminary action ensures that hardware changes are completed proactively, minimizing or eliminating the time penalty that would otherwise be incurred during the failover process itself
Data Source
AI summary
Violation of a recovery requirement is prevented at the time of a failover to a secondary site in case of failure while suppressing cost on hardware of the secondary site in normal times. In a computer system including: a primary site storage system; a secondary site storage system; and a management device, the management device can change a resource of the secondary site storage system, and in the case where a failure occurs in the primary site, the management device performs a failover of making a corresponding secondary volume take over operation of the primary volume, controls so as to enhance a resource of the secondary site storage system, and controls the failover so that a secondary volume which starts operating by the failover before the enhancement of the resource and a secondary volume which starts operating by the failover after the enhancement of the resource exist.


