Unstructured Document Sensitivity Scanning and Risk Assessment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer security methods fail to effectively identify and assess the sensitivity of information in unstructured electronic documents, particularly in unstructured data where the sensitive nature is not trivially visible, leading to potential data breaches and compliance issues across multiple jurisdictions.

Innovation Solution

A method that scans electronic documents and their metadata using machine learning algorithms to classify sensitive data, determine risk scores, and compute exposure risk scores, utilizing cryptographic hashes and knowledge bases to identify and evaluate sensitive information in unstructured documents, and assess compliance risks across different regulatory frameworks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional scanning methods are used to identify sensitive data in unstructured documents, then the scanning process is simple and fast, but the sensitivity detection accuracy is low and sensitive data cannot be effectively identified

Engineering Contradiction:
Improvesensitivity detection accuracyVSAvoidscanning system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary classification system between traditional scanning and sensitive data identification. Machine learning classifiers act as mediators that analyze document content, metadata, and contextual information to determine sensitivity levels, thereby improving detection accuracy without requiring direct complex analysis of all data points

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The scanning system is segmented into multiple independent components: content analysis module, metadata analysis module, classification module, and risk scoring module. Each component handles specific tasks independently, improving overall accuracy while allowing parallel processing that mitigates the complexity burden

Inventive Principle:
Principle #1Segmentation

2Reliability

If comprehensive scanning of all electronic documents is performed to assess sensitivity, then complete risk assessment is achieved, but the processing time and computational resources increase significantly

Engineering Contradiction:
Improverisk assessment completenessVSAvoiddocument processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs partial scanning by initially analyzing only metadata and document properties to identify high-risk documents. Full content analysis is then applied only to documents that exceed certain risk thresholds, achieving comprehensive risk assessment for critical documents while reducing overall processing time through selective deep analysis

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs preliminary classification based on metadata, file type, and document structure before conducting full content analysis. This preliminary action filters out low-risk documents early in the process, allowing comprehensive assessment of only those documents that require detailed examination

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If machine learning algorithms are used to classify sensitive data in unstructured documents, then classification accuracy improves, but the computational complexity and processing overhead increase

Engineering Contradiction:
Improvedata classification accuracyVSAvoidcomputational energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system dynamically adjusts classification parameters such as model complexity, analysis depth, and confidence thresholds based on document characteristics. For simple documents, lighter models with lower computational requirements are used, while complex documents trigger more rigorous analysis, optimizing energy consumption across the document population

Inventive Principle:
Principle #35Parameter changes

4Measurement precision

If multiple risk scores are calculated for each sensitive data occurrence, then the exposure risk assessment becomes more accurate, but the computational load and processing time increase

Engineering Contradiction:
Improveexposure risk scoring accuracyVSAvoiddocument processing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

Different risk scoring methods are applied to different portions of documents based on their sensitivity characteristics. High-risk data elements receive comprehensive multi-factor scoring analysis, while lower-risk elements receive simplified scoring, maintaining overall assessment accuracy while preserving processing throughput

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11188657B2Method and system for managing electronic documents based on sensitivity of information
Publication Date: 2021.11.30 NETGOVERN INC
  • US11188657B2 patent drawing
  • US11188657B2 patent drawing

AI summary

There is described a method for determining a level of sensitivity of information in an electronic document. The method comprises scanning a computer location to select the electronic document, such as an unstructured document in which the sensitive nature of a given portion of the contents is not trivial. In the electronic document, contents and metadata of the electronic document are scanned, and each occurrence of sensitive data is identified by classifying each portion of the contents forming the electronic document as sensitive, or not sensitive, per se. For each occurrence of the sensitive data, there are determined a type of the sensitive data and a risk score associated to the type of the sensitive data, for example from a knowledge base. Using the risk score of each occurrence of the sensitive data, one can determine an exposure risk score of the electronic document.