Machine Learning File Classification for Automated Governance Remediation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Evolving data security and compliance risks in decentralized and complex storage systems, particularly in traditional on-premises and non-integrated architectures, complicate the effective governance and management of digital items, including the identification and permissioning of sensitive information.
Innovation Solution
A machine learning-informed method for automated digital file handling, utilizing classification models to curate sub-corpora of digital files based on type and sensitivity, and executing workflows to modify storage residency, access permissions, and metadata, with configurable policies and heuristics for intelligent data governance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional on-premises decentralized storage systems are used, then data storage flexibility is maintained, but data security and compliance governance become difficult
Solution Approach 1:
The patent introduces a centralized data governance platform as an intermediary layer between decentralized storage systems and compliance requirements. This platform provides classification, tagging, and policy enforcement capabilities that enable secure governance of distributed data without requiring changes to the underlying storage architecture.
Solution Approach 2:
The patent replaces manual data governance processes with automated machine learning-based classification systems. The ML models automatically classify data, assign sensitivity tags, and enforce policies, substituting human-operated mechanical processes with intelligent automated systems that improve both security and operational ease.
2Reliability
If manual data classification and permissioning processes are used, then data governance control is maintained, but processing efficiency and productivity decrease
Solution Approach 1:
The patent enables data to self-classify through automated machine learning processes. The system automatically analyzes data content, determines sensitivity levels, assigns appropriate tags and permissions, and enforces governance policies without human intervention, thereby maintaining control while dramatically improving processing efficiency.
Solution Approach 2:
The patent transforms static manual classification processes into dynamic automated systems that continuously analyze and reclassify data based on changing conditions. The system monitors data access patterns, content changes, and policy updates, automatically adjusting classifications and permissions to maintain governance control at scale.
3Reliability
If centralized cloud-based storage architectures are adopted, then data governance and security management are improved, but system complexity and migration difficulty increase
Solution Approach 1:
The patent segments data governance functionality into independent modular components including data classification, sensitivity tagging, policy enforcement, and access control. These modular segments can be implemented incrementally and work with both centralized and decentralized storage architectures, reducing overall system complexity and migration difficulty.
4Reliability
If comprehensive data classification and automated workflows are implemented, then data security and compliance are enhanced, but computational resources and processing time increase
Solution Approach 1:
The patent applies partial classification actions by focusing ML analysis on critical data elements and metadata rather than processing entire data sets uniformly. The system classifies data at appropriate granularity levels and applies governance policies selectively, reducing computational overhead while maintaining security and compliance effectiveness.
Data Source
AI summary
Systems and methods for automated digital filing includes classifying digital files and curating a number of corpora of digital files based on the classifications associated with each of the digital files. In response to the classification label associated with each distinct corpus of digital files, selectively executing a set of computer instructions associated with the classification label that, when executed, modifies a digital state of a target corpus of digital files.


