Privacy Management Platform Using Confidence Levels for Personal Information Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data protection systems struggle to identify and classify personal information across an organization's data centers, failing to determine the identity of data subjects and find contextual personal information, leading to inefficiencies in managing data risk and customer privacy.

Innovation Solution

A privacy management platform using machine learning classifiers to scan systems, filter out false positives, and correlate true positives to specific data subjects, providing an inventory of personal information indexed by attribute, with features like data risk scoring and natural language queries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If machine learning models are used to determine confidence levels for personal information classification, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces confidence levels as an intermediary metric between the machine learning classification process and the final personal information identification. This intermediary allows the system to quantify uncertainty and make more accurate classifications by comparing confidence levels against thresholds, thereby improving measurement precision while managing the complexity of the ML through a structured evaluation framework

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameter of classification from binary (personal information or not) to a confidence level spectrum (0-1). This parameter transformation enables more nuanced classification decisions and improves measurement precision by allowing the system to express degrees of certainty, while the confidence threshold mechanism manages the complexity of interpreting ML outputs

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If machine learning models analyze multiple features for field-to-field comparisons, then manufacturing precision is improved, but loss of time increases

Engineering Contradiction:
Improveclassification precisionVSAvoidsearch time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-processing and analyzing features of data source fields before the actual personal information search. The system extracts and stores key features (data types, formats, patterns) in advance, so that during the search phase, the machine learning model can quickly compare these pre-analyzed features against target fields, reducing search time while maintaining high classification precision through multi-feature comparison

Inventive Principle:
Principle #10Preliminary action

3Reliability

If comprehensive scanning of all data sources is performed, then reliability is improved, but productivity decreases

Engineering Contradiction:
Improvedata protection reliabilityVSAvoidscan speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies local quality by focusing the comprehensive scan on specific data source fields that have been identified as potentially containing personal information. Rather than uniformly scanning all fields with equal intensity, the system uses machine learning to identify high-probability fields and concentrates scanning resources there, thereby maintaining reliability for critical areas while improving overall productivity by avoiding exhaustive scanning of low-risk fields

Inventive Principle:
Principle #3Local quality

4Measurement precision

If machine learning models compare multiple data sources, then measurement precision is improved, but device complexity increases

Engineering Contradiction:
Improveidentification accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex task of cross-source comparison by breaking it down into field-to-field comparisons between identity data sources and scanned data sources. The machine learning model processes comparisons in discrete, manageable units (individual field pairs), and the system organizes results by data source, field, and confidence level. This segmentation approach improves identification accuracy through systematic multi-source validation while managing processing complexity through structured, modular comparison operations

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3837615B1Machine learning system and methods for determining confidence levels of personal information findings
Publication Date: 2026.03.25 BIGID INC
  • EP3837615B1 patent drawingFigure 1
  • EP3837615B1 patent drawingFigure 2
  • EP3837615B1 patent drawingFigure 3

AI summary

Privacy management platforms are disclosed herein to scan any number of data sources in order to provide users with visibility into stored personal information, risk associated with storing such information and/or usage activity relating to such information. The platforms may correlate personal information findings to specific data subjects and may employ machine learning models to classify findings as corresponding to a particular personal information attribute to provide an indexed inventory across multiple data sources.