Log File Processing for Orphaned Data Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Enterprise systems face challenges in managing and tracking data sets due to frequent changes, such as old devices being taken offline and new applications being added, leading to orphaned data that is difficult to identify and manage effectively.
Innovation Solution
A method and system that processes log files to determine the location of data files within a data environment, allowing for operations such as secure erasure, data migration, security compliance, and efficiency enhancements by filtering and processing log files to identify and manage data operation events.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If replica copies of data are generated and stored at remote locations to meet business continuation and data recovery demands, then data protection and recovery capability are improved, but device complexity and data management difficulty increase due to orphaned data from offline devices and applications
Solution Approach 1:
The system continuously monitors data operation events through log files that record data movements, creations, and deletions. This feedback mechanism enables the cataloging system to track the current state of data and identify orphaned data sets by comparing recorded operations with actual data locations, thereby managing complexity while maintaining protection capabilities.
Solution Approach 2:
A data catalog acts as an intermediary between the complex data environment and management systems. The catalog stores metadata about data files, their locations, and operational history, simplifying the management of replica copies and orphaned data by providing a centralized view without requiring direct management of all data components.
2Duration of action of stationary object
If data sets are maintained for archival purposes along with real-time copies, then data preservation is improved, but difficulty of detecting and measuring orphaned data increases due to lack of visibility into data existence
Solution Approach 1:
The system performs preliminary cataloging of data files by processing log files that record data operation events before orphaned data becomes problematic. By continuously tracking and cataloging data locations and operations in advance, the system maintains visibility into data existence and can identify orphaned data sets before they cause issues.
Solution Approach 2:
The cataloging system uses feedback from log file processing to continuously update the data catalog with current data locations and operational status. This feedback loop ensures that the system maintains accurate information about data existence and can detect orphaned data by comparing cataloged information with current system state.
3Adaptability or versatility
If applications are frequently added and removed from the system to adapt to changing business needs, then system adaptability is improved, but data management difficulty increases due to orphaned data from removed applications
Solution Approach 1:
The system monitors data operation events through continuous log file processing, providing feedback about data locations and operations. This enables the system to automatically detect when data becomes orphaned due to application removals and manage it accordingly, maintaining ease of operation despite frequent application changes.
Solution Approach 2:
The cataloging system performs self-service by automatically processing log files, identifying orphaned data, and managing data sets without requiring manual intervention. The system autonomously tracks data operations, detects orphaned data, and can execute management actions, reducing the operational burden on administrators during frequent application changes.
Data Source
AI summary
A method, computer program product, and computing system for includes processing a log file to determine the location of one or more data files within a data environment. The log file indicates the occurrence of a data operation event within the data environment. A data operation is performed on at least a portion of the one or more data files located via the log file.


