Orphan File Detection in Data Migration Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data migration systems face inefficiencies due to the creation of missing parent and orphan files, which disrupt the relationships between primary and secondary storage devices, leading to inconsistencies and resource wastage.

Innovation Solution

A policy engine server manages the migration of files between primary and secondary storage devices, identifying and updating placeholder files and secondary files to maintain valid references, thereby locating and eliminating missing parent and orphan files through a method that includes scanning for file identification data and updating references in both location and content addressable storage schemes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data migration is performed between primary and secondary storage devices, then storage efficiency is improved, but file relationship consistency deteriorates due to creation of missing parent and orphan files

Engineering Contradiction:
Improvestorage efficiencyVSAvoidfile relationship consistency
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system performs preliminary actions by scanning for file identification data before migration operations complete, identifying missing parent and orphan files in advance, and updating references proactively to maintain consistency throughout the migration process

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback mechanisms by continuously scanning for file identification data, detecting inconsistencies in file relationships, and updating references based on detected issues to ensure ongoing consistency between primary and secondary storage devices

Inventive Principle:
Principle #23Feedback

2Ease of operation

If placeholder files are created during data migration, then file access continuity is maintained, but system complexity increases due to additional file management requirements

Engineering Contradiction:
Improvefile access continuityVSAvoidfile management complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically scanning for file identification data, detecting orphan and missing parent files, and updating references without requiring manual intervention, thereby maintaining file access continuity while managing complexity autonomously

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The scanning and updating mechanisms serve multiple functions: identifying orphan files, detecting missing parent files, updating references, and maintaining consistency across both location and content addressable storage schemes, reducing the need for separate specialized processes

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If reference updates are performed frequently to maintain consistency, then file relationship accuracy is improved, but processing time increases

Engineering Contradiction:
Improvefile relationship accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs periodic scanning for file identification data at appropriate intervals during migration operations, updating references in batches rather than continuously, which maintains accuracy while reducing processing time through scheduled rather than continuous operation

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS7685177B1Detecting and managing orphan files between primary and secondary data stores
Publication Date: 2010.03.23 EMC IP HLDG CO LLC
  • US7685177B1 patent drawing
  • US7685177B1 patent drawing
  • US7685177B1 patent drawing

AI summary

A method and system for locating and eliminating orphan files within a secondary storage device. The method includes identifying a secondary file on a secondary storage device, the secondary file being associated with file identification data, identifying a placeholder file on a primary storage device, the placeholder file being associated with an offline reference, and determining if the offline reference of the placeholder file validly references the secondary file.