Semantic Data Categorization Using Embeddings and LLM Labels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data classification methods require manual effort and are time-consuming, especially in environments with high volumes of data, and rules-based approaches suffer from inaccuracies due to the inability to understand semantic meaning.
Innovation Solution
Utilizing machine learning models to generate embedding vectors for uncategorized data items, clustering them based on similarity, and applying classification labels using large language models to automate data categorization, thereby enabling efficient and accurate data management policy application.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual classification rules are generated to categorize data items, then classification accuracy can be improved, but the time consumption and difficulty increase significantly
Solution Approach 1:
The patent replaces manual mechanical classification processes with machine learning models that automatically generate embedding vectors and perform clustering. This substitution eliminates the need for manual rule creation while achieving accurate classification through semantic understanding of data content.
Solution Approach 2:
The system enables self-service classification by allowing data items to be automatically categorized through ML-generated embeddings and clustering algorithms. The classification system serves itself without human intervention, continuously improving through automated label generation and policy application.
2Ease of manufacture
If rules-based approaches are used for data classification, then the implementation is simple, but the accuracy deteriorates due to inability to understand semantic meaning
Solution Approach 1:
The patent transforms the classification approach by changing from rule-based parameters to semantic parameters. ML models generate embedding vectors that capture semantic meaning, and clustering algorithms group data based on semantic similarity rather than rigid rule matching, fundamentally changing how classification is performed.
Solution Approach 2:
The patent introduces embedding vectors as an intermediary between raw data and classification labels. These vectors serve as a bridge that enables semantic understanding, allowing the system to comprehend the meaning of data content rather than merely matching patterns or keywords.
3Productivity
If automated classification using ML models is implemented, then productivity and accuracy are improved, but the system complexity increases
Solution Approach 1:
The patent segments the classification system into distinct functional modules: data processing module, embedding generation module, clustering module, and label generation module. This segmentation allows each component to be optimized independently and simplifies the overall system architecture by dividing complex tasks into manageable parts.
Solution Approach 2:
The patent creates a universal classification framework that can handle multiple data types and classification scenarios through a single ML-based system. The embedding and clustering mechanism serves multiple functions including semantic understanding, similarity measurement, and automatic clustering, reducing the need for separate specialized systems.
4Reliability
If manual policy application is performed for each data item, then policy accuracy can be ensured, but the time loss increases with huge volumes of data
Solution Approach 1:
The patent applies classification labels to data items immediately after clustering, before policy application is needed. This preliminary classification action enables policies to be pre-associated with label types, so when data items are classified, the corresponding policies are automatically retrieved and applied without delay, ensuring both accuracy and efficiency.
Data Source
AI summary
Broadly speaking, the present techniques provide an automatic way of classifying data items within an environment (e.g. a business, workplace, organisation, etc.). This is advantageous over existing techniques which require manual classification of data items, which is time consuming in environments where hundreds of new data items may be generated in a day or week. The present techniques use an embedding machine learning, ML, model and an LLM to automatically determine the relevant classification label(s) for an unlabelled data item.


