ML File Classification Routing for Data Governance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data handling and governance systems face challenges in efficiently classifying and managing sensitive data due to complex storage architectures, misclassified digital items, and unstructured data, leading to difficulties in data security and compliance.
Innovation Solution
A machine learning-based method for accelerated content classification and routing of digital files, which involves identifying digital files, routing them through a service-defined model instantiation and execution sequence, and using machine learning-based filename classification models to compute content classification inferences and execute instructions for appropriate data handling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional on-premises data storage and nonintegrated storage architectures are used, then data storage flexibility is maintained, but data security and compliance management become difficult
Solution Approach 1:
The system segments data classification into multiple stages using an ensemble of machine learning models, where each model handles specific aspects of classification. This divides the complex security management task into manageable components that can be independently optimized and executed in sequence.
Solution Approach 2:
The patent introduces machine learning-based classification models as intermediary components between raw data and storage systems. These models act as mediators that automatically analyze data characteristics and determine appropriate storage locations, eliminating the need for complex manual security management architectures.
2Measurement precision
If multiple machine learning models are used for content classification, then classification accuracy is improved, but computation time increases
Solution Approach 1:
The system performs preliminary actions by executing simpler classification models first to obtain initial classification results. These preliminary results are then used to inform subsequent more complex model executions, reducing their computational burden and overall processing time while maintaining high accuracy.
Solution Approach 2:
The patent implements partial action by using an ensemble approach where not all models need to execute fully for every data item. The system can stop after achieving sufficient classification confidence through partial model execution, reducing computation time while maintaining acceptable accuracy levels.
3Measurement precision
If comprehensive data classification is performed on all digital files, then data governance accuracy is improved, but processing speed decreases
Solution Approach 1:
The system segments the data processing workflow into distinct phases: initial rapid classification using lightweight models, followed by detailed classification of only those items requiring higher accuracy. This segmentation allows most data to be processed quickly while ensuring comprehensive classification only where necessary.
Solution Approach 2:
The patent applies local quality by differentiating classification intensity based on data characteristics. High-priority or sensitive data items receive comprehensive classification analysis, while routine data items receive streamlined classification, optimizing the balance between accuracy and processing speed for different data types.
Data Source
AI summary
A system and method for accelerated content classification and routing of digital files in a data handling and data governance service includes identifying a digital computer file; sequentially routing the digital computer file to one or more machine learning-based content classification models of a plurality of distinct machine learning-based content classification models based on a service-defined model instantiation and execution sequence, wherein: the service-defined model instantiation and execution sequence defines a model instantiation and execution order for the plurality of distinct machine learning-based content classification models that enables a fast content classification of the digital computer file while minimizing a computation time or runtime of the one or more machine learning-based content classification models; computing, via a machine learning-based filename classification model, a content classification inference based on extracted filename feature data of the digital computer file; and executing one or more computer-executable instructions based on the content classification inference.


