Unstructured Document Sensitivity Scanning and Risk Assessment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer security methods fail to effectively identify and assess the sensitivity of information in unstructured electronic documents, particularly in unstructured data where the sensitive nature is not trivially visible, leading to potential data breaches and compliance issues across multiple jurisdictions.
Innovation Solution
A method that scans electronic documents and their metadata using machine learning algorithms to classify sensitive data, determine risk scores, and compute exposure risk scores, utilizing cryptographic hashes and knowledge bases to identify and evaluate sensitive information in unstructured documents, and assess compliance risks across different regulatory frameworks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional scanning methods are used to identify sensitive data in unstructured documents, then the scanning process is simple and fast, but the sensitivity detection accuracy is low and sensitive data cannot be effectively identified
Solution Approach 1:
The patent introduces an intermediary classification system between traditional scanning and sensitive data identification. Machine learning classifiers act as mediators that analyze document content, metadata, and contextual information to determine sensitivity levels, thereby improving detection accuracy without requiring direct complex analysis of all data points
Solution Approach 2:
The scanning system is segmented into multiple independent components: content analysis module, metadata analysis module, classification module, and risk scoring module. Each component handles specific tasks independently, improving overall accuracy while allowing parallel processing that mitigates the complexity burden
2Reliability
If comprehensive scanning of all electronic documents is performed to assess sensitivity, then complete risk assessment is achieved, but the processing time and computational resources increase significantly
Solution Approach 1:
The system performs partial scanning by initially analyzing only metadata and document properties to identify high-risk documents. Full content analysis is then applied only to documents that exceed certain risk thresholds, achieving comprehensive risk assessment for critical documents while reducing overall processing time through selective deep analysis
Solution Approach 2:
The system performs preliminary classification based on metadata, file type, and document structure before conducting full content analysis. This preliminary action filters out low-risk documents early in the process, allowing comprehensive assessment of only those documents that require detailed examination
3Measurement precision
If machine learning algorithms are used to classify sensitive data in unstructured documents, then classification accuracy improves, but the computational complexity and processing overhead increase
Solution Approach 1:
The system dynamically adjusts classification parameters such as model complexity, analysis depth, and confidence thresholds based on document characteristics. For simple documents, lighter models with lower computational requirements are used, while complex documents trigger more rigorous analysis, optimizing energy consumption across the document population
4Measurement precision
If multiple risk scores are calculated for each sensitive data occurrence, then the exposure risk assessment becomes more accurate, but the computational load and processing time increase
Solution Approach 1:
Different risk scoring methods are applied to different portions of documents based on their sensitivity characteristics. High-risk data elements receive comprehensive multi-factor scoring analysis, while lower-risk elements receive simplified scoring, maintaining overall assessment accuracy while preserving processing throughput
Data Source
AI summary
There is described a method for determining a level of sensitivity of information in an electronic document. The method comprises scanning a computer location to select the electronic document, such as an unstructured document in which the sensitive nature of a given portion of the contents is not trivial. In the electronic document, contents and metadata of the electronic document are scanned, and each occurrence of sensitive data is identified by classifying each portion of the contents forming the electronic document as sensitive, or not sensitive, per se. For each occurrence of the sensitive data, there are determined a type of the sensitive data and a risk score associated to the type of the sensitive data, for example from a knowledge base. Using the risk score of each occurrence of the sensitive data, one can determine an exposure risk score of the electronic document.

