File Tagging for Data Leakage Prevention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for controlling computer files, such as digital rights management and file encryption, often result in false-positives and false-negatives, interfering with normal intra-company communication and failing to reliably distinguish confidential from non-confidential data.
Innovation Solution
A system that scans data intended for external communication, identifies the absence of a tag associated with the data, and interrupts the communication process, using a unified threat management facility to enforce corporate policies through data tagging, allowing flexible control of electronic data transfer without disrupting normal operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional file control methods (DRM, encryption, content scanning) are used to control computer files, then data security is improved, but false-positives and false-negatives increase and normal intra-company communication is interfered with
Solution Approach 1:
The system applies tags to data in advance before communication occurs. Tags are attached to data during creation or storage, so that when communication is attempted, the data is already identified and classified, eliminating the need for real-time content scanning and allowing seamless enforcement of communication policies.
Solution Approach 2:
The invention extracts the control mechanism from content-based scanning to tag-based identification. Instead of scanning file contents during communication attempts, the system uses separate tags that are attached to data, effectively separating the identification function from the communication process and eliminating false-positives and false-negatives.
2Measurement precision
If content scanning methods are used to discriminate confidential data, then data identification capability is improved, but reliability decreases due to false-positives and false-negatives
Solution Approach 1:
The system introduces tags as an intermediary element between data and control mechanisms. Instead of directly scanning and analyzing data contents, the system uses tags as mediators that carry classification information. This intermediary approach eliminates the reliability issues of content scanning while maintaining precise data identification.
Solution Approach 2:
The invention creates a simplified copy or representation of data classification information through tags. Rather than analyzing the actual data content, the system uses tag copies that contain the essential classification attributes, providing reliable and consistent identification without the errors of direct content analysis.
3Reliability
If strict file control policies are enforced to prevent data leakage, then data security is improved, but productivity decreases due to communication disruptions
Solution Approach 1:
Communication policies are enforced based on pre-applied tags rather than real-time content analysis. Since tags are already attached to data before communication attempts, the system can immediately determine whether communication should be allowed, blocked, or require approval, eliminating disruptions to normal communication workflows while maintaining security.
Solution Approach 2:
Data carries its own classification information through tags, enabling self-identification and self-classification. This allows the communication control system to make decisions based on inherent data properties rather than requiring external scanning or analysis, streamlining the enforcement process and maintaining communication efficiency.
Data Source
AI summary
In embodiments of the present invention improved capabilities are described for providing data protection through the detection of tags associated with data or a file. In embodiments the present invention may provide for a step A, where data may be scanned that is intended to be communicated from the client computing facility. In response to step A, at step B, restricted data may be identified by identifying an absence of a tag associated with the data. And finally, in response to step B, at step C, an interruption to the intended communication may be caused.


