Semantic Document Classification for Sensitive Data Access Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems struggle to accurately identify and manage sensitive data in unstructured documents due to reliance on pattern matching and user-driven classification, leading to false positives, computational inefficiencies, and ineffective data security.
Innovation Solution
A method and electronic device that utilize semantic categorization to embed custom metadata within documents, determining privacy personas and classification labels based on semantic categories, enabling precise detection and management of sensitive information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Difficulty of detecting and measuring
If pattern matching and regular expressions are used to identify sensitive data, then detection capability is improved, but false positives increase and computational cost increases
Solution Approach 1:
The patent replaces mechanical pattern matching systems with semantic understanding systems. Instead of using regular expressions and keyword patterns to detect sensitive data, the system employs large language models to comprehend the contextual meaning of document contents, enabling accurate identification of sensitive information while reducing false positives.
Solution Approach 2:
The patent changes the detection parameter from surface-level pattern matching to deep semantic understanding. By transforming the detection approach from checking for specific digit patterns or keywords to analyzing the meaning and context of data, the system achieves higher detection accuracy without being constrained by predefined patterns.
2Reliability
If comprehensive pattern matching is performed on all documents, then detection coverage is improved, but computational overhead increases
Solution Approach 1:
The patent replaces computationally intensive mechanical pattern matching across all documents with a more efficient semantic analysis system. Large language models process document meanings in context, achieving comprehensive detection coverage without the quadratic computational overhead of exhaustive pattern matching on every document.
3Adaptability or versatility
If user-driven classification is implemented, then classification flexibility is improved, but classification accuracy deteriorates
Solution Approach 1:
The patent replaces manual user-driven classification with automated semantic classification powered by large language models. The system analyzes document meanings and contexts to automatically assign appropriate classification labels, maintaining flexibility in classification schemes while eliminating human error and inconsistency.
Solution Approach 2:
The system enables documents to classify themselves through semantic analysis. Large language models automatically evaluate document contents and assign classification labels without requiring user intervention, making the classification process both accurate and self-service oriented.
4Measurement precision
If context words are used to reduce false positives, then detection precision is improved, but system complexity increases
Solution Approach 1:
The patent replaces the complex mechanical system of context word analysis with a unified semantic understanding system. Instead of checking for specific surrounding keywords to validate pattern matches, large language models naturally comprehend contextual relationships, simplifying the system architecture while maintaining or improving detection precision.
Data Source
AI summary
Embodiments herein disclose a method for managing at least one document by an electronic device. The method includes receiving at least one document in an electronic form, where the document includes a plurality of content. Further, the method includes determining at least one semantic category associated with the document and a metadata associated with the document. Further, the method includes embedding custom data within the at least one document, where the custom data comprises the at least one of the semantic category associated with the at least one document. Further, the method includes determining at least one among privacy persona and classification label of the document based on the semantic category associated with the document and the metadata associated with the document. Further, the method includes determining access to the document based on at least one among the privacy persona and the classification label of the document.


