Metadata-Based Personal Information Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing solutions are inefficient in scanning unstructured data stores for personal information due to their large size and complexity, making it difficult to identify and protect sensitive data within unstructured content like documents, which is common in today's business environment.
Innovation Solution
The use of machine learning models that analyze document metadata to predict the presence of personal information, reducing the need to scan entire document contents and thus minimizing computation time and memory requirements, by preprocessing metadata to generate features that are used to determine the likelihood of personal information presence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional scanning algorithms are used to scan unstructured data stores, then complete content analysis can be performed, but scan time becomes excessively long and computation resources are overwhelmed
Solution Approach 1:
The patent segments the document into multiple regions of interest based on metadata indicators, rather than scanning the entire document content. This segmentation allows the system to focus computational resources only on relevant sections, significantly reducing scan time while maintaining detection accuracy for personal information.
Solution Approach 2:
The system performs preliminary analysis by first examining document metadata (such as file name, document type, and other attributes) before scanning the actual content. This preliminary action identifies documents that are likely to contain personal information, allowing the system to prioritize or skip certain documents, thereby reducing overall scan time while maintaining high detection accuracy.
2Productivity
If sampling combined with intelligent correlation algorithms is used, then computation time is reduced, but the methodology is not suitable for unstructured data due to its inherent complexity
Solution Approach 1:
The patent changes the parameters used for analysis by focusing on specific metadata attributes and document characteristics rather than analyzing complete document content. This parameter change enables efficient processing of unstructured data while maintaining reliable identification of personal information through targeted analysis of key features.
3Productivity
If machine learning models analyze only document metadata, then computation time and memory requirements are reduced, but the ability to accurately detect personal information may be compromised
Solution Approach 1:
The patent transitions from analyzing the traditional content dimension to analyzing the metadata dimension. By examining document attributes, file names, document types, and other metadata features, the system achieves accurate prediction of personal information presence without the computational burden of full content analysis, effectively adding a new dimension of analysis that is both efficient and accurate.
Data Source
AI summary
Systems, methods and apparatuses are disclosed to efficiently and accurately scan a plurality of documents located in any number of unstructured data sources. Preprocessed metadata is generated for each document and metadata features are determined based on the preprocessed metadata. A trained machine learning system may utilize the metadata features to predict whether each of the documents contains personal information, without requiring any information relating to the content of such documents.


