Column Classification Models for PHI Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing classification-based predictive data analysis solutions face efficiency and reliability challenges when processing named data collections, particularly in identifying Protected Health Information (PHI) and performing column classification and anomaly detection, due to limitations in handling varied data formats and missing metadata.
Innovation Solution
The implementation of machine-learning-based techniques that generate varied training data through column augmentation and sub-column sampling, allowing column classification machine learning models to perform both classification and anomaly detection efficiently, and a deep learning framework for automatic PHI identification and data anonymization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing classification-based predictive data analysis solutions are used to process named data collections, then basic classification functionality is provided, but efficiency and reliability are insufficient when handling varied data formats and missing metadata
Solution Approach 1:
The patent changes the parameters of the machine learning model by generating multiple sample sequences with different sample sizes from the same training column. This allows the model to learn robust classification patterns that are invariant to variations in data format and length, thereby improving reliability while maintaining adaptability to varied data formats.
Solution Approach 2:
The patent performs preliminary data preparation by generating multiple sample sequences and permuting them before training the machine learning model. This preliminary action ensures that the model is exposed to diverse data configurations in advance, enabling it to handle varied data formats and missing metadata reliably during inference.
2Reliability
If machine learning models are trained with multiple sample sequences and permutations to improve robustness, then reliability and anomaly detection capability are enhanced, but training time and computational resources increase
Solution Approach 1:
The patent applies partial action by generating multiple sample sequences with different sample sizes rather than using all possible permutations. This selective approach provides sufficient diversity for robust training while avoiding the excessive computational burden of exhaustive permutation, thereby reducing training time while maintaining reliability.
3Reliability
If column augmentation and sub-column sampling are used to generate varied training data, then model robustness and anomaly detection performance improve, but data processing complexity increases
Solution Approach 1:
The patent segments the training column into multiple sample sequences with different sample sizes. This segmentation creates diverse training inputs that improve model robustness while maintaining manageable processing complexity through systematic subsampling rather than complex data transformation.
4Productivity
If manual PHI identification and data anonymization processes are used, then data privacy protection is achieved, but processing efficiency and accuracy are reduced
Solution Approach 1:
The patent replaces manual PHI identification and anonymization processes with a machine learning-based automated system. The trained column classification model automatically identifies PHI columns and triggers anonymization, substituting mechanical manual processes with intelligent automated processing that improves both efficiency and accuracy.
Solution Approach 2:
The system enables self-service by allowing the machine learning model to automatically identify PHI columns and initiate anonymization without human intervention. The model serves itself by detecting sensitive data patterns and triggering the appropriate privacy protection actions autonomously.
Data Source
AI summary
There is a need for more effective and efficient performing classification-based predictive data analysis on named data collections. This need can be addressed by, for example, solutions for performing classification-based predictive data analysis on named data collections that utilize at least one of techniques for generating column classification machine learning models to perform column classification, techniques for generating column classification machine learning models to perform anomaly detection, techniques for utilizing trained column classification machine learning models to perform column classification, and techniques for utilizing trained column classification machine learning models to perform anomaly detection.


