Automated Asset Sensitivity Inference for Metadata-Based DLP
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for determining the sensitivity of assets in an organization require labor-intensive human intervention and are prone to errors, and in some cases, manual inspection is not feasible due to the presence of intellectual property information.
Innovation Solution
A system and method that automatically infers asset sensitivity using file system metadata and user activities, employing machine learning models to generate a sensitivity score and DLP alerts without manual tagging or content inspection, and implements DLP policies based on user risk levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual inspection methods are used to determine asset sensitivity, then human judgment can identify sensitive information, but the process becomes labor-intensive and error-prone
Solution Approach 1:
The patent replaces manual human inspection with automated machine learning models that analyze file system metadata and user activities. The system uses algorithms to infer asset sensitivity scores without human intervention, substituting mechanical human judgment with automated computational analysis of behavioral patterns and metadata characteristics.
Solution Approach 2:
The system enables assets to self-classify by automatically generating sensitivity scores based on their own metadata and associated user activities. The machine learning models process asset-specific information autonomously, allowing the system to self-determine sensitivity levels without requiring external human classification efforts.
2Loss of information
If manual tagging is performed to classify assets, then sensitivity information can be recorded, but the process requires significant human resources and time
Solution Approach 1:
The system performs preliminary classification by continuously analyzing file system metadata and user activities in the background. Sensitivity scores are generated proactively before any manual inspection would be needed, with the machine learning models constantly updating asset classifications based on observed behaviors and metadata changes.
Solution Approach 2:
Manual tagging operations are completely replaced by automated machine learning inference. The system uses algorithms to determine sensitivity scores based on patterns in user activities and metadata, eliminating the need for human taggers while maintaining comprehensive sensitivity information capture.
3Measurement precision
If content inspection is performed to determine sensitivity, then accurate classification can be achieved, but intellectual property protection becomes compromised
Solution Approach 1:
The system extracts and analyzes only the necessary metadata characteristics and user activity patterns needed for sensitivity inference, without extracting or inspecting the actual content of protected assets. By taking out only the behavioral and contextual information needed for classification, the system achieves accurate sensitivity determination while leaving the actual intellectual property content untouched and protected.
Solution Approach 2:
The machine learning model acts as an intermediary that infers sensitivity from indirect observations of user behaviors and metadata patterns rather than directly examining asset content. This intermediary approach allows the system to determine sensitivity scores without direct contact with or exposure to the protected intellectual property information.
4Productivity
If automated systems are implemented for sensitivity inference, then productivity increases and human error decreases, but system complexity increases
Solution Approach 1:
The machine learning model serves multiple functions simultaneously: it analyzes diverse metadata types, processes various user activity patterns, generates sensitivity scores, and updates classifications continuously. This multi-functional approach consolidates what would otherwise require multiple separate systems into a single unified platform, managing complexity while maintaining high productivity.
Data Source
AI summary
A method for implementing data loss prevention (DLP) includes: generating an asset lineage map from file system metadata; identifying, based on the asset lineage map, an input feature linked to the asset, a type of the asset, and a plurality of activities linked to the asset; obtaining a sensitivity score for the asset based on the input feature and the type of the asset; obtaining, based on the plurality of activities, a malicious score and a data loss score for the asset; determining a user level of a user; and initiating implementation of a first DLP policy for the user based on the user level, the malicious score, the data loss score, and the sensitivity score.


