Backup Agent Using Re-entrant Child ISRs for Abort Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data backup and restore systems face interruptions due to unexpected events like power outages or network failures, requiring restarts from the beginning and potentially violating service level agreements by failing to meet time windows for data transfer, which can lead to data corruption during restore operations.

Innovation Solution

Implementing a method that invokes a parent interrupt service routine to poll for abort events, saves the backup state using re-entrant child ISRs, and resumes the backup from the saved state after the event is remedied, allowing uninterrupted data transfer and restore operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data transfer is restarted from the beginning after an abort event, then the system can ensure data consistency, but the time required to complete the backup increases and service level agreements may be violated

Engineering Contradiction:
Improvedata consistencyVSAvoidbackup completion time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary actions by continuously monitoring for abort events during the data transfer process and saving checkpoint information at regular intervals. When an abort event occurs, the system has already prepared checkpoint data that allows resumption from the last valid state rather than restarting from the beginning, thus maintaining data consistency while reducing time loss.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If data transfer is restarted from the beginning after an abort event, then data consistency can be maintained, but the number of operations required increases and productivity decreases

Engineering Contradiction:
Improvedata consistencyVSAvoidbackup operation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent enables continuity of useful action by implementing a checkpoint mechanism that allows the backup operation to resume from the last saved state after an abort event. Instead of interrupting the overall backup process and restarting from the beginning, the system maintains continuity by picking up where it left off, thus preserving productivity while ensuring data consistency through the checkpoint validation process.

Inventive Principle:
Principle #20Continuity of useful action

3Reliability

If multiple abort events occur during a data transfer window, then the system can handle individual failures, but the cumulative time loss causes service level agreement violations

Engineering Contradiction:
Improvefailure handling capabilityVSAvoidSLA compliance time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent prepares for multiple abort events by continuously saving checkpoint information during the data transfer process. When multiple abort events occur, the system has pre-saved checkpoint data from each interval, allowing it to quickly resume from the most recent valid checkpoint rather than restarting the entire backup process each time, thus handling failures reliably while minimizing cumulative time loss and maintaining SLA compliance.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If the restore process is interrupted by an abort event, then the system can detect the failure, but data corruption occurs when the restore fails due to the unexpected abort event

Engineering Contradiction:
Improvefailure detectionVSAvoiddata corruption
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent converts the harmful effect of abort events into a beneficial outcome by implementing a checkpoint and validation mechanism. When an abort event occurs during restore, the system detects the failure and uses the saved checkpoint information to determine the last valid state. This allows the system to roll back to the checkpoint state, transforming the potential harm of data corruption into a controlled recovery process that actually improves data integrity by eliminating corrupted data.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS20200409799A1Stream level uninterrupted backup operation using an interrupt service routine approach
Publication Date: 2020.12.31 EMC IP HLDG CO LLC
  • US20200409799A1 patent drawing
  • US20200409799A1 patent drawing
  • US20200409799A1 patent drawing

AI summary

Embodiments are described for performing an uninterrupted backup in a storage system in view of one or more abort events. A backup agent receives writes one or more data blocks to a write latch. A parent interrupt service routine (ISR) polls for abort events. In response to an abort event, an intermediate interrupt is generated that spawns a child processes for each process of the backup. The intermediate ISR logs each child ISR, the process it is responsible for, and the intermediate interrupt, for later restoration of the backup state. After a recovery of the above event, then each child ISR can be called to restore its state. After restoring the state, the backup agent resumes the backup from where the abort event was detected. The child ISRs are re-entrant. If another abort event is detected, the backup state can again be saved and later resumed from that state.