Incremental Backup Hard Links via Inode Change Logs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing backup systems face performance degradation and resource wastage due to the need to process unnecessary disk accesses and frequent full backups caused by conventional change logs, especially in large file systems with high change rates, and the challenge of efficiently tracking hardlinks during incremental backups.

Innovation Solution

The implementation of a system that uses a change log with unique identifiers, such as inode numbers, to index file changes, and efficiently manages hardlinks by creating new and deleted file name lists to avoid unnecessary disk accesses and limit the change log size, allowing for incremental backups based on these lists.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a conventional change log records every change in temporal order for all files, then complete file change history is maintained, but backup performance degrades due to processing unnecessary disk accesses and the change log size becomes unmanageably large

Engineering Contradiction:
Improvecomplete file change historyVSAvoidbackup performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the change log processing by file type, separating hardlinks from regular files. Hardlinks are tracked using an inode-based mechanism that groups multiple filenames pointing to the same data, while regular files use traditional change log records. This segmentation allows the system to process only relevant changes during backup operations, eliminating unnecessary disk accesses for deleted files while maintaining complete change history.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary inode-based tracking mechanism that sits between the file system and the change log. Instead of directly recording every file change in temporal order, the system uses inodes as intermediaries to aggregate hardlink information and determine which files actually need backup. This intermediary layer filters out redundant change log records while preserving the complete history of file modifications.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the change log grows to record all file changes over time, then complete backup history is maintained, but the change log size becomes too large requiring frequent purging or wrapping

Engineering Contradiction:
Improvebackup history completenessVSAvoidchange log size
Core Design Contradiction:
ReliabilityVSVolume of stationary object

Solution Approach 1:

The patent segments change log management by file type, maintaining separate tracking mechanisms for hardlinks (using inodes) and regular files. This segmentation prevents the change log from growing indefinitely by only recording relevant changes and using compact inode-based storage for hardlink information, reducing the overall volume of stored data while preserving complete backup history.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter of change log storage by transitioning from temporal ordering of all file changes to an inode-based indexing system. This parameter change allows the system to store change information in a more compact form, grouping related changes by inode and eliminating redundant records, thereby maintaining complete backup history with reduced storage requirements.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If a backup application trawls the entire file system to generate a list of changed files, then all changed files are identified, but significant computing resources are consumed

Engineering Contradiction:
Improvechanged file identification accuracyVSAvoidcomputing resource consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent introduces an intermediary change log and inode tracking mechanism that pre-processes and organizes file change information. Instead of trawling the entire file system during backup operations, the system queries the pre-populated change log and inode structures to identify changed files. This intermediary mechanism maintains precise identification of changed files while dramatically reducing computing resource consumption during backup execution.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Loss of time

If incremental backups are performed using conventional change logs, then backup time is reduced, but performance degrades due to processing unnecessary disk accesses for deleted files

Engineering Contradiction:
Improvebackup timeVSAvoidbackup performance
Core Design Contradiction:
Loss of timeVSProductivity

Solution Approach 1:

The patent segments the incremental backup process by file type, using inode-based tracking for hardlinks to determine which files actually exist and need to be backed up. This segmentation eliminates the processing of unnecessary disk accesses for deleted files that plague conventional change log approaches, maintaining short backup times while preserving backup performance through targeted file selection.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10204016B1Incrementally backing up file system hard links based on change logs
Publication Date: 2019.02.12 EMC IP HLDG CO LLC
  • US10204016B1 patent drawing
  • US10204016B1 patent drawing
  • US10204016B1 patent drawing

AI summary

Incrementally backing up file system hard links based on change logs is described. A system identifies a unique identifier and a file name associated with a file event in a file system. The system determines whether a change log for the file system lacks an association between the unique identifier and the file name. The system adds the file name to one of a new file name list and a deleted file name list associated with the change log in response to a determination that the change log for the file system lacks the association between the unique identifier and the file name. The system incrementally backs up the file event based on at least one of the new file name list and the deleted file name list.