Float Feature Vector Clustering for Sensitive Data Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data protection systems face challenges in accurately identifying sensitive data due to primitive classification methods, high false positives and negatives, inability to handle large volumes of files, and require significant human resources for rule management, especially in large corporate environments.
Innovation Solution
A system that calculates float feature vectors for files, clusters them, generates DNA vectors for each cluster, and compares these vectors to identify sensitive data, using machine learning algorithms and context data for enhanced accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If statistical fingerprinting technologies are used to generate digital fingerprints and compare them against a database, then the ability to identify sensitive data is improved, but the system generates a large number of false positives and false negatives
Solution Approach 1:
The patent changes the parameter space from traditional fingerprint matching to high-dimensional float vector embeddings. By transforming file characteristics into dense vector representations and using similarity search in this new parameter space, the system achieves more accurate and reliable sensitive data identification with fewer false positives and negatives.
Solution Approach 2:
The patent replaces the mechanical fingerprint database comparison system with a machine learning-based vector embedding system. Instead of relying on pre-defined fingerprint patterns, the system uses neural networks to learn representations of file characteristics, enabling more intelligent and accurate classification.
2Ease of manufacture
If primitive classification methods using rule engines are used to scan file contents, then the system is easier to implement, but the ability to accurately identify sensitive data deteriorates
Solution Approach 1:
The patent replaces the mechanical rule-engine based classification system with a machine learning-based vector embedding system. Instead of relying on hand-crafted rules, the system uses neural networks to automatically learn representations of file characteristics from training data, achieving superior accuracy while maintaining ease of deployment through automated model training.
Solution Approach 2:
The system performs self-service by automatically training its own models and generating vector embeddings without requiring manual rule configuration. The machine learning models automatically adapt to different file types and sensitive data patterns, eliminating the need for analysts to define robust rule sets.
3Measurement precision
If statistical fingerprinting methods are used to monitor files, then the detection capability is improved, but the system cannot handle a large number of files due to database size constraints
Solution Approach 1:
The patent extracts the essential characteristics of files into compact vector representations, separating the important information from the bulk of file data. This extraction allows the system to work with millions of files by operating on the compressed vector representations rather than the full file contents, dramatically increasing scalability.
Solution Approach 2:
The patent transforms the data representation from raw file contents to high-dimensional float vector embeddings. This parameter transformation enables efficient storage and comparison of millions of files by working with compact vector representations, overcoming the database size constraints of traditional fingerprint methods.
4Measurement precision
If traditional DP systems are deployed, then the ability to identify sensitive data is improved, but the amount of human resources required for managing rules and policies increases
Solution Approach 1:
The patent implements self-service by using machine learning models that automatically learn and adapt to sensitive data patterns without human intervention. The models are trained on labeled data and can independently perform classification tasks, eliminating the need for analysts to manually manage and update rules and policies.
Solution Approach 2:
The patent replaces the mechanical rule management system with an automated machine learning system. Instead of requiring human experts to define, maintain, and update classification rules, the system uses neural networks that automatically learn from training data and can be retrained as needed, dramatically reducing the human resource burden.
Data Source
AI summary
Methods and systems for identifying sensitive data (SD) stored on data repositories is disclosed. The data is processed to calculate a plurality of float feature (FF) vectors associated with the data. The FF vectors are clustered into a plurality of clusters, each cluster associated with a respective subset of the data. A DNA vector representative of the cluster is generated for each cluster. The DNA vectors of respective clusters are compared to one or more FF vectors calculated for a respective one or more user supplied examples of SD. One or more clusters are classified as SD based on the result of the comparing, thereby identifying respective subsets of data as SD.


