Privacy Management Platform Using Confidence Levels for Personal Information Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data protection systems struggle to identify and classify personal information across an organization's data centers, failing to determine the identity of data subjects and find contextual personal information, leading to inefficiencies in managing data risk and customer privacy.
Innovation Solution
A privacy management platform using machine learning classifiers to scan systems, filter out false positives, and correlate true positives to specific data subjects, providing an inventory of personal information indexed by attribute, with features like data risk scoring and natural language queries.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If machine learning models are used to determine confidence levels for personal information classification, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent introduces confidence levels as an intermediary metric between the machine learning classification process and the final personal information identification. This intermediary allows the system to quantify uncertainty and make more accurate classifications by comparing confidence levels against thresholds, thereby improving measurement precision while managing the complexity of the ML through a structured evaluation framework
Solution Approach 2:
The system changes the parameter of classification from binary (personal information or not) to a confidence level spectrum (0-1). This parameter transformation enables more nuanced classification decisions and improves measurement precision by allowing the system to express degrees of certainty, while the confidence threshold mechanism manages the complexity of interpreting ML outputs
2Manufacturing precision
If machine learning models analyze multiple features for field-to-field comparisons, then manufacturing precision is improved, but loss of time increases
Solution Approach 1:
The patent performs preliminary actions by pre-processing and analyzing features of data source fields before the actual personal information search. The system extracts and stores key features (data types, formats, patterns) in advance, so that during the search phase, the machine learning model can quickly compare these pre-analyzed features against target fields, reducing search time while maintaining high classification precision through multi-feature comparison
3Reliability
If comprehensive scanning of all data sources is performed, then reliability is improved, but productivity decreases
Solution Approach 1:
The patent applies local quality by focusing the comprehensive scan on specific data source fields that have been identified as potentially containing personal information. Rather than uniformly scanning all fields with equal intensity, the system uses machine learning to identify high-probability fields and concentrates scanning resources there, thereby maintaining reliability for critical areas while improving overall productivity by avoiding exhaustive scanning of low-risk fields
4Measurement precision
If machine learning models compare multiple data sources, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent segments the complex task of cross-source comparison by breaking it down into field-to-field comparisons between identity data sources and scanned data sources. The machine learning model processes comparisons in discrete, manageable units (individual field pairs), and the system organizes results by data source, field, and confidence level. This segmentation approach improves identification accuracy through systematic multi-source validation while managing processing complexity through structured, modular comparison operations
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Privacy management platforms are disclosed herein to scan any number of data sources in order to provide users with visibility into stored personal information, risk associated with storing such information and/or usage activity relating to such information. The platforms may correlate personal information findings to specific data subjects and may employ machine learning models to classify findings as corresponding to a particular personal information attribute to provide an indexed inventory across multiple data sources.