Evolving Data Classification for Safe Stale File Remediation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data remediation processes face challenges in accurately classifying and managing data, leading to potential disruptions in business operations and legal/regulatory risks due to indiscriminate data deletion.
Innovation Solution
An evolving classification model that iteratively trains and refines its classification abilities over time, placing unclassified data in quarantine for later evaluation, and applies data retention policies based on improved classification confidence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is removed to reduce storage costs and improve efficiency, then storage costs decrease and workflow efficiency improves, but business operations may be impeded and legal/regulatory risks increase
Solution Approach 1:
The system performs preliminary classification of data before deletion using an evolving AI model. Data is evaluated and categorized in advance, with uncertain or potentially important data placed in quarantine rather than immediately deleted, preventing disruption to business operations while still reducing storage costs through systematic removal of confirmed ROT data
Solution Approach 2:
The classification model operates with feedback mechanisms that continuously learn from results and adjust its classifications. The model is retrained over time based on organizational data patterns, improving accuracy in distinguishing between deletable ROT data and data that should be retained for business operations, thereby reducing false deletions while maintaining storage efficiency
2Quantity of substance
If data is removed to eliminate redundant and obsolete data, then storage costs decrease and data management improves, but legal and regulatory implications arise
Solution Approach 1:
The system performs preliminary classification of data before deletion using an evolving AI model. Data is evaluated and categorized in advance, with uncertain or potentially important data placed in quarantine rather than immediately deleted, preventing disruption to business operations while still reducing storage costs through systematic removal of confirmed ROT data
Solution Approach 2:
The classification model operates with feedback mechanisms that continuously learn from results and adjust its classifications. The model is retrained over time based on organizational data patterns, improving accuracy in distinguishing between deletable ROT data and data that should be retained for business operations, thereby reducing false deletions while maintaining storage efficiency
3Measurement precision
If a classification model is used to identify data for deletion, then data management accuracy improves, but the model requires continuous training and refinement over time
Solution Approach 1:
The classification model performs self-improving through automated retraining processes. The system automatically collects feedback from classification results and uses it to refine the model's parameters, enabling the model to improve its own accuracy over time without requiring manual intervention or complex external training infrastructure
Solution Approach 2:
The model training and refinement process operates continuously rather than as a one-time event. The system maintains ongoing learning from organizational data patterns, ensuring the classification accuracy remains high and adapts to changing business requirements without interrupting data management operations
Data Source
AI summary
This disclosure describes techniques for performing data remediation. In one example, this disclosure describes a method that includes identifying a plurality of stale files; applying a classification model to each of the plurality of stale files; identifying a plurality of unclassified files, wherein each of the unclassified files is one of the plurality of stale files that the classification model was not able to classify with a confidence level that exceeds a threshold confidence level; updating the classification model, over a period of time, to generate an evolved classification model; applying the evolved classification model to each of the unclassified files; identifying a subset of the unclassified files that the evolved classification model was not able to classify with a confidence level that exceeds the threshold confidence level; and deleting each of the files in the subset of the unclassified files.


