PII Identification System Using Standard Data Bank
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems lack an efficient method for identifying personally identifiable information (PII) within large datasets, which is crucial for protecting sensitive data and complying with regulations like HIPAA, as they often require manual processing and are not scalable for extensive data models.
Innovation Solution
A system and method that utilize a processor to identify PII within data models by applying processing rules, comparing identified data with a standard database, validating matches, and marking validated PII, while also detecting and adding differing PII to the database, enabling efficient scanning of both physical and logical data models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual processing methods are used to identify PII, then accuracy can be maintained through human judgment, but productivity is significantly reduced and the process is not scalable for extensive data models
Solution Approach 1:
The system segments the PII identification process into distinct functional modules: data ingestion component, processing rules engine, standard data bank, validation module, and marking system. Each module handles a specific aspect of PII identification, allowing the complex task to be distributed across multiple specialized components that can process data in parallel, thereby increasing throughput without overwhelming system complexity
Solution Approach 2:
The patent introduces a standard data bank as an intermediary layer between the processing rules and the final PII identification. This standard data bank stores pre-defined PII patterns, formats, and validation criteria, acting as a mediator that translates complex identification logic into searchable, reusable standards. This intermediary structure enables automated processing while maintaining accuracy through established benchmarks
2Productivity
If automated processing rules are applied to identify PII, then productivity increases and scalability is improved, but measurement precision may deteriorate due to rule-based limitations
Solution Approach 1:
The system performs preliminary actions by pre-compiling comprehensive processing rules and populating the standard data bank with established PII patterns before the actual identification process begins. This preliminary preparation includes categorizing different types of PII (personal names, addresses, phone numbers, social security numbers) with their specific formats and validation criteria, enabling the automated system to quickly match incoming data against pre-established standards without sacrificing accuracy
Solution Approach 2:
The patent implements a feedback mechanism where the system continuously validates identified PII against the standard data bank and adjusts processing rules based on validation results. When the marking system identifies potential PII, it cross-references with the standard data bank, and discrepancies feed back into refining the processing rules. This closed-loop feedback ensures that automated processing maintains high precision by learning from and adapting to validation outcomes
3Measurement precision
If comprehensive processing rules are implemented to cover all PII types, then measurement precision is improved, but device complexity increases making the system harder to maintain
Solution Approach 1:
The standard data bank is designed as a universal repository that handles multiple types of PII through a single standardized interface. Instead of creating separate processing systems for different PII categories (names, addresses, phone numbers, etc.), the patent implements a multi-functional standard data bank that can store and validate various PII formats using consistent data structures and validation methods. This universal approach maintains comprehensive coverage while reducing overall system complexity through standardization
4Reliability
If the system validates each identified PII against the standard data bank, then reliability of PII identification is improved, but loss of time increases due to additional validation steps
Solution Approach 1:
The system applies partial validation action by performing confidence-based validation rather than exhaustive verification of every identified PII. The marking system uses the standard data bank to generate confidence scores for each identified PII match, and only performs full validation steps for high-value or low-confidence matches. This partial action approach maintains high reliability for critical PII while reducing overall validation time by applying lighter validation to high-confidence matches
Data Source
AI summary
The system may be configured to perform operations including identifying, by a processor, personally identifiable information (PII) within a data model based on processing rules, to create identified PII, wherein the data model comprises entity information about an entity; comparing the identified PII with established PII in a standard data bank; validating the identified PII in response to the identified PII matching the established PII, to create validated PII; and marking the validated PII with a PII marker in response to the validating the identified PII.


