ML Data Classification for Confidentiality Anomaly Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Data Loss Prevention (DLP) solutions face challenges in accurate automatic confidentiality classification of unstructured data due to insufficient labelled data, requiring manual classification and being ineffective in adapting to the ever-changing threat landscape.

Innovation Solution

A method and system for content and context-aware data classification using machine learning, which extracts features from documents into TF-IDF and LSI vectors, classifies documents into confidentiality categories, and includes a machine learning engine with model training and evaluation modules to predict category classifications, reducing the need for manual data classification and improving adaptability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual data confidentiality classification is used, then classification accuracy can be maintained, but time consumption and labor requirements increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service classification by automatically processing unstructured data through machine learning models. The automated confidentiality classification system performs classification tasks independently without requiring manual intervention for each document, thereby maintaining accuracy while significantly reducing time consumption and labor requirements.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual classification process with an automated machine learning-based system. The machine learning model processes documents algorithmically, substituting human labor with computational mechanisms that can classify large volumes of data quickly and consistently, thus resolving the contradiction between accuracy and time consumption.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If supervised machine learning methods are used, then classification accuracy can be improved, but the amount of labelled data required increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidamount of labelled data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system performs preliminary actions by pre-processing and extracting features from unlabelled data before classification. Feature extraction and representation learning are conducted in advance on the entire dataset, enabling the model to learn patterns from unlabelled data and reduce dependency on labelled data for achieving accurate classification.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies parameter changes by transforming the classification approach from traditional supervised learning to a hybrid model that incorporates unsupervised learning techniques. This changes the fundamental parameters of the learning process, allowing the system to achieve high accuracy with reduced labelled data by leveraging both labelled and unlabelled data through advanced machine learning algorithms.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If existing DLP solutions are implemented, then basic data protection can be provided, but adaptability to changing threat landscape is insufficient

Engineering Contradiction:
Improvedata protection capabilityVSAvoidadaptability to threat landscape
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system introduces dynamics by implementing continuous learning and adaptation capabilities. The machine learning model continuously processes new data and updates its classifications, enabling the system to adapt to evolving threat landscapes dynamically. This dynamic approach allows the DLP solution to maintain reliability while improving its adaptability to changing security challenges.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent incorporates feedback mechanisms where classification results and security outcomes are fed back into the machine learning model for continuous improvement. This feedback loop enables the system to learn from actual security events and adjust its future classifications, thereby enhancing both the reliability of data protection and the adaptability to emerging threats through iterative optimization.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12033040B2Method, machine learning engines and file management platform systems for content and context aware data classification and security anomaly detection
Publication Date: 2024.07.09 DATHENA SCI PTE LTD
  • US12033040B2 patent drawing
  • US12033040B2 patent drawing
  • US12033040B2 patent drawing

AI summary

Systems, methods and computer readable medium are provided for perform a method for content and context aware data classification or a method for content and context aware data security anomaly detection. The method for content and context aware data confidentiality classification includes scanning one or more documents in one or more network data repositories of a computer network and extracting content features and context features of the one or more documents into one or more term frequency-inverse document frequency (TF-IDF) vectors and one or more latent semantic indexing (LSI) vectors. The method further includes classifying the one or more documents into a number of category classifications by machine learning the extracted content features and context features of the one or more documents at a file management platform of the computer network, each of the category classifications being associated with one or more confidentiality classifications.