Data Retention Scoring for Predictive Analytics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Big data analytics systems face challenges in balancing storage costs and retaining data that could provide insights into IT infrastructure performance, as purged data may contain key information for predictive analytics and preventing downtime, leading to a dilemma between longer retention periods and increased costs or reduced historical data availability.

Innovation Solution

A system comprising an identify unit, a score unit, and a select unit that identifies, scores, and selectively retains data above a threshold, using different schemes to measure relevancy for predicting system behavior, allowing only critical data to be stored for extended periods, thereby optimizing costs and enhancing predictive analytics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is retained for longer periods, then predictive analytics capability is improved, but storage costs increase

Engineering Contradiction:
Improvepredictive analytics capabilityVSAvoidstorage costs
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies local quality by differentiating data retention policies based on data type and value. Instead of uniformly retaining all data, the system identifies and retains only high-value data (metrics, events, logs) that contribute to predictive analytics, while allowing lower-value data to be purged. This selective retention approach maintains predictive capability while reducing storage costs.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts retention parameters based on data characteristics. By changing the retention parameter from a fixed time-based policy to a value-based policy (using data scoring and classification), the system optimizes the balance between predictive analytics capability and storage costs.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If all data is retained, then historical data availability is improved, but storage costs increase

Engineering Contradiction:
Improvehistorical data availabilityVSAvoidstorage costs
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent extracts and retains only the essential historical data that provides key insights into IT infrastructure performance. The system purges non-essential data while maintaining availability of critical historical information needed for analytics, thereby reducing storage costs without significant loss of valuable information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system discards low-value data that does not contribute to predictive analytics, while recovering and retaining high-value data. This selective discarding and recovering process optimizes storage resource allocation by maintaining only the data that provides meaningful historical insights.

Inventive Principle:
Principle #34Discarding and recovering

3Quantity of substance

If data is purged to reduce costs, then storage costs are optimized, but predictive analytics capability deteriorates

Engineering Contradiction:
Improvestorage costsVSAvoidpredictive analytics capability
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system applies local quality by implementing differentiated retention policies for different data types (metrics, events, logs) based on their specific value to predictive analytics. High-value data is retained while low-value data is purged, maintaining analytics capability while optimizing storage costs.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system changes the retention parameter from time-based to value-based, using data scoring mechanisms to determine which data to retain. This parameter change ensures that data purging does not deteriorate predictive analytics capability, as only data below a certain value threshold is discarded.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10503766B2Retain data above threshold
Publication Date: 2019.12.10 HEWLETT PACKARD ENTERPRISE DEV LP
  • US10503766B2 patent drawing
  • US10503766B2 patent drawing

AI summary

According to an example, different types of data stored at a database may be identified. The identified data may be scored, where different types of data are scored according to different schemes. The scored data that is above a threshold may be selectively retained. The different schemes may relate to measuring a relevancy of the identified data for predicting behavior of a system.