Utility-Based Timeseries Data Purging Algorithm
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data purging methods, such as time-based purging, result in significant loss of business intelligence data by ignoring relationships between monitoring data samples, leading to abrupt loss of knowledge and compromised analysis capabilities.
Innovation Solution
A purging algorithm that assigns 'utility values' to data samples based on their importance, using models that capture relationships between timeseries, regions of interest, and age of data, to minimize information loss while preserving high-value samples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If time-based purging is used to reduce storage cost and management overhead, then storage cost and data management cost are reduced, but business intelligence information is lost abruptly and relationships between data samples are ignored
Solution Approach 1:
The patent changes the purging parameter from simple time-based deletion to utility-value-based selection. Each data sample is assigned a utility value reflecting its importance to business intelligence, and purging decisions are made based on these values rather than uniformly by time threshold. This resolves the contradiction by preserving high-utility samples even if they are old, while deleting low-utility samples regardless of recency.
Solution Approach 2:
The patent performs preliminary assignment of utility values to data samples before purging. By pre-evaluating and tagging each sample's importance level, the system prepares the data repository for intelligent purging that maintains business intelligence value. This preliminary classification enables the system to selectively preserve critical data relationships while reducing storage costs.
2Quantity of substance
If data samples are purged to meet repository capacity limits, then storage capacity constraints are satisfied, but data analysis capability is compromised
Solution Approach 1:
The patent applies different retention policies to different data samples based on their local utility characteristics. Rather than uniform purging, each sample is evaluated individually for its contribution to analysis capability. High-utility samples that are critical for maintaining analysis reliability are preserved, while low-utility samples are purged to meet capacity constraints. This localized quality approach ensures capacity management without compromising overall analysis capability.
3Ease of manufacture
If all data samples before a threshold time are deleted to simplify purging implementation, then ease of implementation is improved, but relationships between cascaded events and their impacts are lost
Solution Approach 1:
The patent introduces feedback mechanisms where the utility value assignment process considers relationships between data samples. The system evaluates how deleting a sample would impact the understanding of cascaded events and their impacts. By incorporating this feedback into the purging decision, the system maintains critical event relationships while still achieving purging objectives, balancing implementation simplicity with information preservation.
Data Source
AI summary
There is disclosed methods, systems and computer program products for purging stored data in a repository. Users attach relative importance to all data samples across all timeseries in a repository. The importance attached to a data sample is the ‘utility value’ of the data sample. An algorithm uses the utility of data samples and allocates the storage space of the repository in such a way that the total loss of information due to purging is minimized while preserving samples with a high utility value.


