Context-Aware Hierarchical Labeling for Sensitive Data Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for detecting and protecting sensitive data in natural language processing environments, particularly in AI systems like LMs and GenAI, suffer from high false positives and negatives, and are unable to scale effectively, compromising data security and privacy.
Innovation Solution
A hierarchical, context-aware labeling mechanism optimized using machine learning techniques applies labels at multiple levels (word, chunk/phrase, document) to ensure precise and scalable protection of sensitive data, leveraging latent semantic structure and reversible tokenization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data protection methods are used in NLP environments, then data security is improved, but detection accuracy deteriorates due to high false positives and negatives
Solution Approach 1:
The patent applies segmentation by implementing a hierarchical labeling system that divides data protection into multiple levels: word-level labels, phrase-level labels, and document-level labels. This segmentation allows the system to detect sensitive information at different granularities, improving detection accuracy by capturing context-dependent sensitivity that single-level approaches miss.
Solution Approach 2:
The patent introduces a new dimension to data protection by adding contextual awareness through hierarchical labeling. Instead of treating all data uniformly, the system adds a labeling dimension that captures semantic context, leading to more accurate detection while maintaining security through multi-level protection strategies.
2Measurement precision
If context-aware hierarchical labeling is implemented, then detection accuracy is improved, but system complexity increases
Solution Approach 1:
The hierarchical labeling system segments the complex task of context-aware detection into manageable levels (word, phrase, document), making the overall system more approachable and easier to implement while maintaining high accuracy through the structured breakdown of complexity.
Solution Approach 2:
The system performs preliminary labeling actions during the data processing pipeline, establishing hierarchical labels in advance before final detection decisions are made. This preliminary structuring of information reduces the computational burden during actual detection operations.
3Measurement precision
If machine learning optimization is applied to the labeling mechanism, then false positives and negatives are minimized, but computational resources increase
Solution Approach 1:
The machine learning model performs preliminary training to learn patterns of sensitive information across different contexts. Once trained, the model can make efficient predictions during actual data processing, reducing the need for computationally intensive real-time analysis while maintaining low false positive and false negative rates.
Solution Approach 2:
The system uses pre-trained machine learning models that capture the essence of sensitive data detection through learned representations. These pre-computed models serve as efficient copies of complex detection logic, enabling accurate false positive reduction without requiring excessive computational resources during execution.
Data Source
AI summary
The present disclosure relates to scalable systems and methods for detecting, labeling, and protecting sensitive data in natural language processing (NLP) environments. This includes NLP applications in artificial intelligence (AI) systems, such as language models (LMs) and generative AI (GenAI). More particularly, the present disclosure introduces a hierarchical, context-aware labeling mechanism that is optimized using an LM in conjunction with machine learning (ML) techniques to ensure the utility-preserving effective protection of sensitive data with, for example, minimal false positives and false negatives and/or optimal precision and recall (e.g., in terms of an F1 Score).


