Automated Asset Sensitivity Inference for Metadata-Based DLP

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for determining the sensitivity of assets in an organization require labor-intensive human intervention and are prone to errors, and in some cases, manual inspection is not feasible due to the presence of intellectual property information.

Innovation Solution

A system and method that automatically infers asset sensitivity using file system metadata and user activities, employing machine learning models to generate a sensitivity score and DLP alerts without manual tagging or content inspection, and implements DLP policies based on user risk levels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual inspection methods are used to determine asset sensitivity, then human judgment can identify sensitive information, but the process becomes labor-intensive and error-prone

Engineering Contradiction:
Improvesensitivity detection accuracyVSAvoidasset inspection efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual human inspection with automated machine learning models that analyze file system metadata and user activities. The system uses algorithms to infer asset sensitivity scores without human intervention, substituting mechanical human judgment with automated computational analysis of behavioral patterns and metadata characteristics.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables assets to self-classify by automatically generating sensitivity scores based on their own metadata and associated user activities. The machine learning models process asset-specific information autonomously, allowing the system to self-determine sensitivity levels without requiring external human classification efforts.

Inventive Principle:
Principle #25Self-service

2Loss of information

If manual tagging is performed to classify assets, then sensitivity information can be recorded, but the process requires significant human resources and time

Engineering Contradiction:
Improvesensitivity information completenessVSAvoidclassification time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary classification by continuously analyzing file system metadata and user activities in the background. Sensitivity scores are generated proactively before any manual inspection would be needed, with the machine learning models constantly updating asset classifications based on observed behaviors and metadata changes.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Manual tagging operations are completely replaced by automated machine learning inference. The system uses algorithms to determine sensitivity scores based on patterns in user activities and metadata, eliminating the need for human taggers while maintaining comprehensive sensitivity information capture.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If content inspection is performed to determine sensitivity, then accurate classification can be achieved, but intellectual property protection becomes compromised

Engineering Contradiction:
Improvesensitivity classification accuracyVSAvoidintellectual property exposure risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system extracts and analyzes only the necessary metadata characteristics and user activity patterns needed for sensitivity inference, without extracting or inspecting the actual content of protected assets. By taking out only the behavioral and contextual information needed for classification, the system achieves accurate sensitivity determination while leaving the actual intellectual property content untouched and protected.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The machine learning model acts as an intermediary that infers sensitivity from indirect observations of user behaviors and metadata patterns rather than directly examining asset content. This intermediary approach allows the system to determine sensitivity scores without direct contact with or exposure to the protected intellectual property information.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Productivity

If automated systems are implemented for sensitivity inference, then productivity increases and human error decreases, but system complexity increases

Engineering Contradiction:
Improveasset classification throughputVSAvoidsystem architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The machine learning model serves multiple functions simultaneously: it analyzes diverse metadata types, processes various user activity patterns, generates sensitivity scores, and updates classifications continuously. This multi-functional approach consolidates what would otherwise require multiple separate systems into a single unified platform, managing complexity while maintaining high productivity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12423428B2Method and system for inferring document sensitivity
Publication Date: 2025.09.23 DTEX SYSTEMS INC
  • US12423428B2 patent drawing
  • US12423428B2 patent drawing
  • US12423428B2 patent drawing

AI summary

A method for implementing data loss prevention (DLP) includes: generating an asset lineage map from file system metadata; identifying, based on the asset lineage map, an input feature linked to the asset, a type of the asset, and a plurality of activities linked to the asset; obtaining a sensitivity score for the asset based on the input feature and the type of the asset; obtaining, based on the plurality of activities, a malicious score and a data loss score for the asset; determining a user level of a user; and initiating implementation of a first DLP policy for the user based on the user level, the malicious score, the data loss score, and the sensitivity score.