Pre-firewall Data Classification Using Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large enterprises face challenges in efficiently classifying and managing personal data at the point of entry or exit to ensure compliance with internal and external regulations, as existing methods are impractical due to time constraints and the complexity of validating data before it is disseminated throughout the enterprise.
Innovation Solution
A system that uses machine-learning processing to classify personal data in-line with or pre-firewall, determining which data owners and data elements require classification, and assigning appropriate security classifications based on learned patterns and metadata, ensuring that data is properly routed and handled according to regulatory standards.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data classification and validation is performed after data files are received and disseminated to repositories and applications, then data security and compliance control is improved, but processing time and operational complexity increase significantly
Solution Approach 1:
The patent applies preliminary action by performing data classification and security validation at the point of data entry or exit (pre-firewall), before data is disseminated throughout the enterprise. This early classification ensures that data is tagged with appropriate security labels and compliance requirements upfront, eliminating the need for time-consuming validation processes later when data has already been distributed to multiple repositories and applications.
Solution Approach 2:
The patent introduces an intermediary classification system that acts as a mediator between data flow and security controls. This intermediary layer classifies data based on predefined criteria and attaches security metadata, enabling downstream systems to automatically handle classified data without requiring complex real-time validation, thus reducing processing time while maintaining security.
2Measurement precision
If comprehensive data validation and classification is performed on all data elements, then compliance accuracy is improved, but processing speed and productivity decrease
Solution Approach 1:
The patent applies local quality by classifying data elements based on their specific characteristics and sensitivity levels rather than treating all data uniformly. The system identifies and applies different classification rules to different data types (e.g., personally identifiable information, financial data, health information), ensuring high compliance accuracy for sensitive data while maintaining faster processing for less sensitive data elements.
Solution Approach 2:
The patent implements partial action by focusing classification efforts on data elements that require compliance handling based on predefined criteria. Rather than validating every single data element comprehensively, the system applies classification selectively to data that matches specific patterns or sensitivity thresholds, achieving sufficient compliance accuracy without the overhead of exhaustive validation of all data.
3Reliability
If data classification is performed at the point of entry/exit (pre-firewall), then data proliferation control is improved, but system complexity increases
Solution Approach 1:
The patent applies universality by designing a classification system that handles multiple data types, compliance requirements, and security policies through a unified pre-firewall classification mechanism. This universal classifier can process various data formats and apply different classification rules based on data content, eliminating the need for separate classification systems for different data types and reducing overall system complexity.
Data Source
AI summary
Classification of personal data in incoming or outgoing data files in-line or pre-firewall. The invention determines which data owners and/or data associated with the data owners requires classification (e.g., which individuals/customers and/or data is applicable to internal or external regulations) and, subsequently determines the classifications and identifies the classifications in the data file the data owners and data within the data file so that the data can be routed according to the identified classifications. In specific embodiments machine-learning processing is used to learn, determine and/or predict which data owners and/or data associated with the individual/customers requires classification and the classifications to assign to those data owners and/or data elements.


