ML Data Classification for Confidentiality Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Data Loss Prevention (DLP) solutions face challenges in accurate automatic confidentiality classification of unstructured data due to insufficient labelled data, requiring manual classification and being ineffective in adapting to the ever-changing threat landscape.
Innovation Solution
A method and system for content and context-aware data classification using machine learning, which extracts features from documents into TF-IDF and LSI vectors, classifies documents into confidentiality categories, and includes a machine learning engine with model training and evaluation modules to predict category classifications, reducing the need for manual data classification and improving adaptability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual data confidentiality classification is used, then classification accuracy can be maintained, but time consumption and labor requirements increase significantly
Solution Approach 1:
The system enables self-service classification by automatically processing unstructured data through machine learning models. The automated confidentiality classification system performs classification tasks independently without requiring manual intervention for each document, thereby maintaining accuracy while significantly reducing time consumption and labor requirements.
Solution Approach 2:
The patent replaces the mechanical manual classification process with an automated machine learning-based system. The machine learning model processes documents algorithmically, substituting human labor with computational mechanisms that can classify large volumes of data quickly and consistently, thus resolving the contradiction between accuracy and time consumption.
2Measurement precision
If supervised machine learning methods are used, then classification accuracy can be improved, but the amount of labelled data required increases
Solution Approach 1:
The system performs preliminary actions by pre-processing and extracting features from unlabelled data before classification. Feature extraction and representation learning are conducted in advance on the entire dataset, enabling the model to learn patterns from unlabelled data and reduce dependency on labelled data for achieving accurate classification.
Solution Approach 2:
The patent applies parameter changes by transforming the classification approach from traditional supervised learning to a hybrid model that incorporates unsupervised learning techniques. This changes the fundamental parameters of the learning process, allowing the system to achieve high accuracy with reduced labelled data by leveraging both labelled and unlabelled data through advanced machine learning algorithms.
3Reliability
If existing DLP solutions are implemented, then basic data protection can be provided, but adaptability to changing threat landscape is insufficient
Solution Approach 1:
The system introduces dynamics by implementing continuous learning and adaptation capabilities. The machine learning model continuously processes new data and updates its classifications, enabling the system to adapt to evolving threat landscapes dynamically. This dynamic approach allows the DLP solution to maintain reliability while improving its adaptability to changing security challenges.
Solution Approach 2:
The patent incorporates feedback mechanisms where classification results and security outcomes are fed back into the machine learning model for continuous improvement. This feedback loop enables the system to learn from actual security events and adjust its future classifications, thereby enhancing both the reliability of data protection and the adaptability to emerging threats through iterative optimization.
Data Source
AI summary
Systems, methods and computer readable medium are provided for perform a method for content and context aware data classification or a method for content and context aware data security anomaly detection. The method for content and context aware data confidentiality classification includes scanning one or more documents in one or more network data repositories of a computer network and extracting content features and context features of the one or more documents into one or more term frequency-inverse document frequency (TF-IDF) vectors and one or more latent semantic indexing (LSI) vectors. The method further includes classifying the one or more documents into a number of category classifications by machine learning the extracted content features and context features of the one or more documents at a file management platform of the computer network, each of the category classifications being associated with one or more confidentiality classifications.


