Float Feature Vector Clustering for Sensitive Data Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data protection systems face challenges in accurately identifying sensitive data due to primitive classification methods, high false positives and negatives, inability to handle large volumes of files, and require significant human resources for rule management, especially in large corporate environments.

Innovation Solution

A system that calculates float feature vectors for files, clusters them, generates DNA vectors for each cluster, and compares these vectors to identify sensitive data, using machine learning algorithms and context data for enhanced accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If statistical fingerprinting technologies are used to generate digital fingerprints and compare them against a database, then the ability to identify sensitive data is improved, but the system generates a large number of false positives and false negatives

Engineering Contradiction:
Improveaccuracy of sensitive data identificationVSAvoidfalse positive and false negative rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent changes the parameter space from traditional fingerprint matching to high-dimensional float vector embeddings. By transforming file characteristics into dense vector representations and using similarity search in this new parameter space, the system achieves more accurate and reliable sensitive data identification with fewer false positives and negatives.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces the mechanical fingerprint database comparison system with a machine learning-based vector embedding system. Instead of relying on pre-defined fingerprint patterns, the system uses neural networks to learn representations of file characteristics, enabling more intelligent and accurate classification.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of manufacture

If primitive classification methods using rule engines are used to scan file contents, then the system is easier to implement, but the ability to accurately identify sensitive data deteriorates

Engineering Contradiction:
Improveease of system implementationVSAvoidaccuracy of sensitive data identification
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent replaces the mechanical rule-engine based classification system with a machine learning-based vector embedding system. Instead of relying on hand-crafted rules, the system uses neural networks to automatically learn representations of file characteristics from training data, achieving superior accuracy while maintaining ease of deployment through automated model training.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system performs self-service by automatically training its own models and generating vector embeddings without requiring manual rule configuration. The machine learning models automatically adapt to different file types and sensitive data patterns, eliminating the need for analysts to define robust rule sets.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If statistical fingerprinting methods are used to monitor files, then the detection capability is improved, but the system cannot handle a large number of files due to database size constraints

Engineering Contradiction:
Improvedetection capabilityVSAvoidnumber of files that can be monitored
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts the essential characteristics of files into compact vector representations, separating the important information from the bulk of file data. This extraction allows the system to work with millions of files by operating on the compressed vector representations rather than the full file contents, dramatically increasing scalability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the data representation from raw file contents to high-dimensional float vector embeddings. This parameter transformation enables efficient storage and comparison of millions of files by working with compact vector representations, overcoming the database size constraints of traditional fingerprint methods.

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If traditional DP systems are deployed, then the ability to identify sensitive data is improved, but the amount of human resources required for managing rules and policies increases

Engineering Contradiction:
Improvesensitive data identification capabilityVSAvoidhuman resources required for rule management
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements self-service by using machine learning models that automatically learn and adapt to sensitive data patterns without human intervention. The models are trained on labeled data and can independently perform classification tasks, eliminating the need for analysts to manually manage and update rules and policies.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical rule management system with an automated machine learning system. Instead of requiring human experts to define, maintain, and update classification rules, the system uses neural networks that automatically learn from training data and can be retrained as needed, dramatically reducing the human resource burden.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11256821B2Method of identifying and tracking sensitive data and system thereof
Publication Date: 2022.02.22 MINEREYE LTD
  • US11256821B2 patent drawing
  • US11256821B2 patent drawing
  • US11256821B2 patent drawing

AI summary

Methods and systems for identifying sensitive data (SD) stored on data repositories is disclosed. The data is processed to calculate a plurality of float feature (FF) vectors associated with the data. The FF vectors are clustered into a plurality of clusters, each cluster associated with a respective subset of the data. A DNA vector representative of the cluster is generated for each cluster. The DNA vectors of respective clusters are compared to one or more FF vectors calculated for a respective one or more user supplied examples of SD. One or more clusters are classified as SD based on the result of the comparing, thereby identifying respective subsets of data as SD.