Non-disruptive Storage Migration in Failover Clusters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data migration in a failover cluster environment is challenging due to the need for fine control over I/O operations and the possibility of aborting and restarting at multiple steps, which requires significant communication and coordination among nodes, leading to inefficient use of system resources, especially for unlikely events like failures.

Innovation Solution

A method for non-disruptively migrating data from a source storage device to a target storage device in a failover cluster, involving metadata creation, synchronization, and a commit operation that ensures seamless transition of read and write operations from the source to the target device, with a roll-forward flag to manage access control and completion of the migration process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If non-disruptive migration is performed with fine control over I/O operations and multiple abort/restart steps, then data integrity and non-disruptive operation are maintained, but communication and coordination overhead among cluster nodes increases significantly

Engineering Contradiction:
Improvedata integrityVSAvoidcoordination overhead
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the coordination and communication overhead into a dedicated migration manager component that operates independently from the regular failover cluster nodes. This migration manager handles all the complex coordination for abort/restart scenarios, I/O suspension control, and state synchronization, separating this overhead from the normal cluster operation paths.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements preliminary actions by establishing a dedicated migration metadata structure and pre-configuring abort/restart handling logic before migration begins. The system pre-sets up the coordination framework, including predefined states and transition rules, so that when abort or restart events occur during migration, the complex coordination work has already been prepared and can execute efficiently without ad-hoc communication overhead.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If I/O operations are temporarily suspended during migration for fine control, then synchronization accuracy is improved, but productivity and system availability decrease

Engineering Contradiction:
Improvesynchronization accuracyVSAvoidsystem availability
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent implements periodic I/O suspension and resumption cycles during migration. Instead of suspending I/O for the entire migration duration, the system periodically suspends I/O operations at controlled intervals to perform synchronization checkpoints, then resumes operations. This allows synchronization accuracy to be maintained through regular checkpoints while minimizing the cumulative impact on system availability.

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The patent applies partial suspension of I/O operations rather than complete suspension. During migration, I/O operations are selectively suspended only when necessary for synchronization checkpoints, while other operations continue. This partial action approach maintains sufficient synchronization accuracy while preserving overall system productivity and availability.

Inventive Principle:
Principle #16Partial or excessive action

3Ease of operation

If access control data is dynamically changed during migration, then seamless transition is achieved, but risk of data loss and system instability increases

Engineering Contradiction:
Improveseamless transitionVSAvoidsystem stability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements feedback mechanisms where access control data changes are monitored and validated at each migration stage. The system continuously checks the state of source and target devices, verifies synchronization status, and only permits access control transitions when predefined conditions are met. This feedback loop detects potential instability conditions and prevents unsafe access control changes, maintaining system reliability during seamless transitions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent prepares cushioning measures in advance by implementing rollback capabilities and pre-validation checks before access control data changes. The system pre-configures fallback options and validates the safety of upcoming transitions, so that if instability occurs during access control changes, the system can revert to the previous safe state, preventing data loss and maintaining system stability.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS8775861B1Non-disruptive storage device migration in failover cluster environment
Publication Date: 2014.07.08 EMC IP HLDG CO LLC
  • US8775861B1 patent drawing
  • US8775861B1 patent drawing
  • US8775861B1 patent drawing

AI summary

A method of performing data migration from a source storage device to a target storage device in a failover cluster includes use of a roll-forward flag to signal successful completion of a migration operation from a migration node to failover nodes of the cluster, reliably controlling host access to the target storage device to ensure that it is used only when it has been successfully synchronized to the source storage device and a commit operation has occurred that ensures that subsequent read and write operations are directed exclusively to the target storage device.