Data Lake Log Analysis for Automated Data Retention Decisions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face challenges in managing large-scale data storage efficiently, including inadequate analysis of data utilization, identification of redundant data, limited insight into data access patterns, and suboptimal data storage optimization, leading to excessive costs, computational burdens, and compliance risks.

Innovation Solution

An improved extract, transform, load (ETL) engine leverages access and inventory logs to perform in-depth analysis of data usage patterns, enabling precise identification of underutilized or obsolete data, and provides granular insights into data access trends, facilitating strategic data lifecycle decisions and improved resource management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data is retained in large-scale data lakes for extended periods, then data availability and compliance requirements are met, but storage costs and computational burdens increase

Engineering Contradiction:
Improvedata availabilityVSAvoidstorage cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system dynamically changes the retention parameters of data by implementing automated data lifecycle management that transitions data between active storage, archival storage, and deletion based on usage patterns and compliance requirements. This resolves the contradiction by adjusting storage duration parameters to match actual needs rather than using fixed retention periods.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system automatically identifies and discards redundant or obsolete data through duplicate detection algorithms and data quality assessment, while recovering valuable data through automated archiving and restoration capabilities. This enables selective retention that reduces storage costs while maintaining necessary data availability.

Inventive Principle:
Principle #34Discarding and recovering

2Loss of information

If comprehensive data access logging is implemented, then data utilization analysis is improved, but system complexity and processing overhead increase

Engineering Contradiction:
Improvedata utilization insightVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system extracts only the most relevant data access information needed for utilization analysis, such as access frequency, data types, and usage patterns, while filtering out redundant logging details. This selective extraction provides sufficient insight into data utilization without the overhead of comprehensive logging of every access event.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system pre-processes and aggregates data access logs in real-time, preparing summary statistics and usage patterns before they need to be analyzed. This preliminary aggregation reduces the complexity of subsequent analysis while maintaining comprehensive utilization insights.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If manual data management processes are used, then data security and compliance can be monitored, but operational efficiency and cost-effectiveness decrease

Engineering Contradiction:
Improvecompliance monitoringVSAvoidoperational efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system implements automated self-service capabilities including automatic data classification, tagging, and compliance tagging based on built-in policies. The system monitors its own data inventory and automatically enforces retention policies without requiring manual intervention, thereby maintaining compliance monitoring while dramatically improving operational efficiency.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system continuously monitors data access patterns and compliance status, then automatically adjusts data management decisions based on this feedback. This closed-loop system ensures compliance requirements are met while optimizing storage and access operations, eliminating the need for manual monitoring processes.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250370974A1Applied programmatic data lake analysis
Publication Date: 2025.12.04 DISNEY ENTERPRISES INC
  • US20250370974A1 patent drawing
  • US20250370974A1 patent drawing
  • US20250370974A1 patent drawing

AI summary

Techniques for identifying and applying data relationships. These techniques include retrieving log data relating to data stored in a plurality of electronic repositories. The techniques further include transforming the log data to generate structured table data, including matching a pattern in the log data to generate the structured table data, and identifying a plurality of data relationships for the data stored in the plurality of electronic repositories. This includes identifying metadata associated with the data stored in a plurality of electronic repositories, and correlating the generated structured table data with the associated metadata. The techniques further include applying the identified plurality of data relationships to at least one of: (i) modify operation of a computer software job operating on the data stored in the plurality of electronic repositories or (ii) identify for removal a portion of the data stored in a plurality of electronic repositories.