Semantic Data Categorization Using Embeddings and LLM Labels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data classification methods require manual effort and are time-consuming, especially in environments with high volumes of data, and rules-based approaches suffer from inaccuracies due to the inability to understand semantic meaning.

Innovation Solution

Utilizing machine learning models to generate embedding vectors for uncategorized data items, clustering them based on similarity, and applying classification labels using large language models to automate data categorization, thereby enabling efficient and accurate data management policy application.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual classification rules are generated to categorize data items, then classification accuracy can be improved, but the time consumption and difficulty increase significantly

Engineering Contradiction:
Improveclassification accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical classification processes with machine learning models that automatically generate embedding vectors and perform clustering. This substitution eliminates the need for manual rule creation while achieving accurate classification through semantic understanding of data content.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service classification by allowing data items to be automatically categorized through ML-generated embeddings and clustering algorithms. The classification system serves itself without human intervention, continuously improving through automated label generation and policy application.

Inventive Principle:
Principle #25Self-service

2Ease of manufacture

If rules-based approaches are used for data classification, then the implementation is simple, but the accuracy deteriorates due to inability to understand semantic meaning

Engineering Contradiction:
Improveimplementation simplicityVSAvoidclassification accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent transforms the classification approach by changing from rule-based parameters to semantic parameters. ML models generate embedding vectors that capture semantic meaning, and clustering algorithms group data based on semantic similarity rather than rigid rule matching, fundamentally changing how classification is performed.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces embedding vectors as an intermediary between raw data and classification labels. These vectors serve as a bridge that enables semantic understanding, allowing the system to comprehend the meaning of data content rather than merely matching patterns or keywords.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated classification using ML models is implemented, then productivity and accuracy are improved, but the system complexity increases

Engineering Contradiction:
Improveclassification speedVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the classification system into distinct functional modules: data processing module, embedding generation module, clustering module, and label generation module. This segmentation allows each component to be optimized independently and simplifies the overall system architecture by dividing complex tasks into manageable parts.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal classification framework that can handle multiple data types and classification scenarios through a single ML-based system. The embedding and clustering mechanism serves multiple functions including semantic understanding, similarity measurement, and automatic clustering, reducing the need for separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If manual policy application is performed for each data item, then policy accuracy can be ensured, but the time loss increases with huge volumes of data

Engineering Contradiction:
Improvepolicy application accuracyVSAvoiddata processing throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies classification labels to data items immediately after clustering, before policy application is needed. This preliminary classification action enables policies to be pre-associated with label types, so when data items are classified, the corresponding policies are automatically retrieved and applied without delay, ensuring both accuracy and efficiency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250384655A1Method for Automatically Categorizing Data Items
Publication Date: 2025.12.18 VARONIS SYSTEMS INC
  • US20250384655A1 patent drawing
  • US20250384655A1 patent drawing
  • US20250384655A1 patent drawing

AI summary

Broadly speaking, the present techniques provide an automatic way of classifying data items within an environment (e.g. a business, workplace, organisation, etc.). This is advantageous over existing techniques which require manual classification of data items, which is time consuming in environments where hundreds of new data items may be generated in a day or week. The present techniques use an embedding machine learning, ML, model and an LLM to automatically determine the relevant classification label(s) for an unlabelled data item.