Automated PII Classification via NLP and ML Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data classification processes for personally identifiable information (PII) rely on manual, resource-intensive rule-based engines that require constant updates and lack efficient access control, making it difficult to balance data availability with sensitive information protection.
Innovation Solution
A system and method using statistical techniques and natural language processing to classify PII data into protection groups based on metadata, enabling automated and granular access control through a centralized, computer-implemented system.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If manual rule-based classification is used, then data classification can be performed, but extensive manual input and substantial resources are required
Solution Approach 1:
The system enables self-service classification by training machine learning models to automatically identify and classify PII data without requiring manual rule updates. The model learns from training data and autonomously classifies new data instances, eliminating the need for continuous manual rule maintenance while improving classification efficiency.
Solution Approach 2:
The patent replaces manual mechanical rule-based classification with automated machine learning-based classification. The system uses trained models to perform classification tasks that previously required human analysts to manually create and update rules, thereby increasing automation and productivity simultaneously.
2Adaptability or versatility
If rule-based engines are used with exact keyword matching, then classification can be performed, but the system needs constant manual updates to include new keywords
Solution Approach 1:
The system performs preliminary action by training machine learning models in advance with comprehensive training data that includes various PII types and patterns. Once trained, the model can adapt to new PII types without requiring manual rule updates, as the pre-trained model generalizes to recognize new patterns automatically.
Solution Approach 2:
The classification system transitions from static keyword-based rules to dynamic machine learning models that can adapt and learn new PII patterns. The model's parameters and decision boundaries are dynamic and can adjust to new data distributions, enabling the system to recognize new PII types without manual intervention.
3Ease of operation
If data is made widely available for business use, then information accessibility improves, but unauthorized access to sensitive information increases
Solution Approach 1:
The system applies local quality by implementing granular access control at the attribute level rather than applying uniform restrictions to all data. Different security policies and protection groups are applied to different PII attributes based on their sensitivity, allowing business stakeholders to access non-sensitive data freely while restricting access to sensitive PII attributes to authorized personnel only.
Data Source
AI summary
An embodiment of the present invention is directed to classifying attributes into respective PI/PG categories based on metadata. An embodiment of the present invention may classify each attribute into PII/Non-PII and then into various Protection group codes that define access, roles permissions, privileges and/or other action. An embodiment of the present invention may leverage various statistical techniques, natural language processing (NLP) methods and different combinations of algorithms customized to improve prediction accuracies of a classifier model.


