Unapplied-Operation Recovery for Synchronous Dataset Replication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing storage systems face inefficiencies in managing data replication and recovery across multiple storage systems, particularly in scenarios where synchronous replication is required, leading to potential data loss and system instability.
Innovation Solution
Implementing a direct-mapped flash storage system where higher-level processes manage data operations directly across flash drives without address translation by lower-level controllers, utilizing non-volatile RAM for quick data buffering, and employing dual storage array controllers for failover and distributed management of storage drives.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If synchronous replication is implemented across storage systems, then data consistency is improved, but system complexity and write operation overhead increase
Solution Approach 1:
The patent segments the storage system into multiple independent storage systems, each capable of autonomous operation. By dividing the replicated data into segments stored across different storage systems, the system achieves data consistency without requiring complex centralized coordination, thus reducing overall system complexity while maintaining reliability.
Solution Approach 2:
The patent implements preliminary actions by pre-configuring replication relationships and recovery mechanisms before failures occur. Recovery configurations are established in advance, allowing the system to quickly restore data consistency after failures without complex real-time decision-making, thereby reducing operational complexity while ensuring reliability.
2Reliability
If synchronous replication is implemented across storage systems, then data consistency is improved, but write operation performance deteriorates
Solution Approach 1:
The patent applies partial action by implementing selective replication where only critical data segments are replicated synchronously across storage systems. Non-critical segments use asynchronous or deferred replication, reducing the performance overhead of write operations while maintaining data consistency for essential data, thus balancing reliability and productivity.
Solution Approach 2:
The patent changes replication parameters dynamically based on data priority, storage system status, and workload conditions. By adjusting replication timing, frequency, and target selection, the system optimizes write performance while ensuring data consistency is maintained for critical operations, resolving the contradiction between reliability and productivity.
3Ease of operation
If recovery actions are delayed until failures occur, then system operation simplicity is improved, but data loss risk increases
Solution Approach 1:
The patent implements preliminary recovery configuration where recovery actions are pre-planned and configured before failures occur. Recovery policies, target selections, and parameter settings are established in advance, allowing the system to automatically execute recovery actions without complex real-time decision-making. This maintains operational simplicity while ensuring rapid response to failures, thereby reducing data loss risk.
Solution Approach 2:
The patent enables self-service recovery where the storage system automatically detects failures and executes pre-configured recovery actions without human intervention. The system monitors its own state, identifies failures, and triggers recovery processes autonomously, maintaining operational simplicity while minimizing data loss through immediate automated response.
Data Source
AI summary
Initiating recovery actions when a dataset ceases to be synchronously replicated across a set of storage systems, including: receiving, by at least one storage system among a plurality of storage systems implementing a symmetric input/output model for a synchronously replicated dataset, a request to modify the dataset; identifying one or more operations associated with the request to modify the dataset that have not been applied to at least one storage system of the plurality of storage systems; and responsive to a system fault among the plurality of storage systems synchronously replicating the dataset, applying a recovery action based on recovery information that identifies one or more operations that have not been applied to the plurality of storage systems.


