Data Retention Scoring for Predictive Analytics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Big data analytics systems face challenges in balancing storage costs and retaining data that could provide insights into IT infrastructure performance, as purged data may contain key information for predictive analytics and preventing downtime, leading to a dilemma between longer retention periods and increased costs or reduced historical data availability.
Innovation Solution
A system comprising an identify unit, a score unit, and a select unit that identifies, scores, and selectively retains data above a threshold, using different schemes to measure relevancy for predicting system behavior, allowing only critical data to be stored for extended periods, thereby optimizing costs and enhancing predictive analytics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is retained for longer periods, then predictive analytics capability is improved, but storage costs increase
Solution Approach 1:
The patent applies local quality by differentiating data retention policies based on data type and value. Instead of uniformly retaining all data, the system identifies and retains only high-value data (metrics, events, logs) that contribute to predictive analytics, while allowing lower-value data to be purged. This selective retention approach maintains predictive capability while reducing storage costs.
Solution Approach 2:
The system dynamically adjusts retention parameters based on data characteristics. By changing the retention parameter from a fixed time-based policy to a value-based policy (using data scoring and classification), the system optimizes the balance between predictive analytics capability and storage costs.
2Loss of information
If all data is retained, then historical data availability is improved, but storage costs increase
Solution Approach 1:
The patent extracts and retains only the essential historical data that provides key insights into IT infrastructure performance. The system purges non-essential data while maintaining availability of critical historical information needed for analytics, thereby reducing storage costs without significant loss of valuable information.
Solution Approach 2:
The system discards low-value data that does not contribute to predictive analytics, while recovering and retaining high-value data. This selective discarding and recovering process optimizes storage resource allocation by maintaining only the data that provides meaningful historical insights.
3Quantity of substance
If data is purged to reduce costs, then storage costs are optimized, but predictive analytics capability deteriorates
Solution Approach 1:
The system applies local quality by implementing differentiated retention policies for different data types (metrics, events, logs) based on their specific value to predictive analytics. High-value data is retained while low-value data is purged, maintaining analytics capability while optimizing storage costs.
Solution Approach 2:
The system changes the retention parameter from time-based to value-based, using data scoring mechanisms to determine which data to retain. This parameter change ensures that data purging does not deteriorate predictive analytics capability, as only data below a certain value threshold is discarded.
Data Source
AI summary
According to an example, different types of data stored at a database may be identified. The identified data may be scored, where different types of data are scored according to different schemes. The scored data that is above a threshold may be selectively retained. The different schemes may relate to measuring a relevancy of the identified data for predicting behavior of a system.

