Time-Based Data Clustering for File System Activity Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing mechanisms for tracking and analyzing user activity in file systems generate vast amounts of data with limited value, requiring further processing to extract useful information.
Innovation Solution
A method that involves associating data-access instances with time-based clusters, iteratively refining their distribution, and facilitating time-density analysis to uncover noteworthy patterns in user data-access activity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data-access tracking mechanisms are implemented to monitor user activity in file systems, then security monitoring and activity analysis capabilities are improved, but the system generates a massive amount of data that provides very little value without further processing
Solution Approach 1:
The patent segments the massive data-access history into meaningful clusters based on temporal patterns. By dividing the continuous stream of data-access instances into discrete time-based clusters, the system transforms overwhelming raw data into manageable, analyzable units that reveal useful patterns while filtering out noise.
Solution Approach 2:
The patent introduces clustering algorithms as an intermediary processing layer between data collection and analysis. This intermediary automatically processes the raw data-access history, identifying and grouping related events, thereby extracting value from the massive dataset without requiring manual intervention or overwhelming storage resources.
2Adaptability or versatility
If comprehensive data-access history is collected for analysis, then the scope and completeness of security monitoring is improved, but the complexity of processing and extracting useful information increases significantly
Solution Approach 1:
The patent performs preliminary clustering of data-access instances before detailed analysis. By pre-grouping events into time-based clusters based on temporal proximity and access patterns, the system reduces the complexity of subsequent analysis while maintaining comprehensive coverage of all data-access activities.
Solution Approach 2:
The patent transforms the analysis approach by changing from examining individual data-access events to analyzing clustered groups of events. This parameter change from event-level to cluster-level analysis significantly reduces processing complexity while preserving the ability to detect security anomalies and usage patterns.
3Quantity of substance
If raw data-access instances are analyzed without clustering, then data completeness is maintained, but the ability to identify useful patterns and insights is significantly reduced
Solution Approach 1:
The patent merges related data-access instances into temporal clusters based on their proximity in time and similarity in access patterns. This combining process preserves all original data points while grouping them into meaningful units that highlight patterns, thereby maintaining data completeness while dramatically improving pattern identification accuracy.
Solution Approach 2:
The patent adds a temporal clustering dimension to the analysis of data-access instances. By organizing events along the time dimension into clusters, the system transforms flat, unordered data into a structured representation that reveals patterns and relationships invisible in the raw data, thereby improving measurement precision without losing data completeness.
Data Source
AI summary
In one embodiment, a method includes accessing a data-access history for a time period, the data-access history comprising a plurality of data-access instances. The method further includes initially associating each data-access instance with a time-based data-access cluster of a plurality of time-based data-access clusters based, at least in part, on a time of the data-access instance. In addition, the method includes iteratively refining a time distribution of the plurality of data-access instances across the plurality of time-based data-access clusters. Further, the method includes facilitating a time-density analysis of the plurality of data-access instances using the iteratively refined plurality of time-based data-access clusters.


