Metadata-Based Personal Information Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing solutions are inefficient in scanning unstructured data stores for personal information due to their large size and complexity, making it difficult to identify and protect sensitive data within unstructured content like documents, which is common in today's business environment.

Innovation Solution

The use of machine learning models that analyze document metadata to predict the presence of personal information, reducing the need to scan entire document contents and thus minimizing computation time and memory requirements, by preprocessing metadata to generate features that are used to determine the likelihood of personal information presence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional scanning algorithms are used to scan unstructured data stores, then complete content analysis can be performed, but scan time becomes excessively long and computation resources are overwhelmed

Engineering Contradiction:
Improvedetection accuracyVSAvoidscan time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the document into multiple regions of interest based on metadata indicators, rather than scanning the entire document content. This segmentation allows the system to focus computational resources only on relevant sections, significantly reducing scan time while maintaining detection accuracy for personal information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary analysis by first examining document metadata (such as file name, document type, and other attributes) before scanning the actual content. This preliminary action identifies documents that are likely to contain personal information, allowing the system to prioritize or skip certain documents, thereby reducing overall scan time while maintaining high detection accuracy.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If sampling combined with intelligent correlation algorithms is used, then computation time is reduced, but the methodology is not suitable for unstructured data due to its inherent complexity

Engineering Contradiction:
Improvescanning efficiencyVSAvoidmapping accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent changes the parameters used for analysis by focusing on specific metadata attributes and document characteristics rather than analyzing complete document content. This parameter change enables efficient processing of unstructured data while maintaining reliable identification of personal information through targeted analysis of key features.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If machine learning models analyze only document metadata, then computation time and memory requirements are reduced, but the ability to accurately detect personal information may be compromised

Engineering Contradiction:
Improveprocessing speedVSAvoidprediction accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent transitions from analyzing the traditional content dimension to analyzing the metadata dimension. By examining document attributes, file names, document types, and other metadata features, the system achieves accurate prediction of personal information presence without the computational burden of full content analysis, effectively adding a new dimension of analysis that is both efficient and accurate.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11100252B1Machine learning systems and methods for predicting personal information using file metadata
Publication Date: 2021.08.24 BIGID INC
  • US11100252B1 patent drawing
  • US11100252B1 patent drawing
  • US11100252B1 patent drawing

AI summary

Systems, methods and apparatuses are disclosed to efficiently and accurately scan a plurality of documents located in any number of unstructured data sources. Preprocessed metadata is generated for each document and metadata features are determined based on the preprocessed metadata. A trained machine learning system may utilize the metadata features to predict whether each of the documents contains personal information, without requiring any information relating to the content of such documents.