Structured Content Data Loss Protection via Flattened Stream Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing user productivity applications face challenges in identifying and protecting sensitive data within structured documents like spreadsheets and presentations, where sensitive information may be split across multiple data entities, making it difficult to prevent data loss and ensure compliance with privacy policies.
Innovation Solution
A data loss protection framework that receives structured user content, transforms it into flattened representations, and uses mapping information to identify sensitive data, marking it within the user interface and providing options for obfuscation, allowing for efficient detection and protection of sensitive content across various applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If sensitive data is distributed across multiple structured data entities (cells, objects, tables), then data flexibility and document complexity are improved, but data loss protection effectiveness deteriorates due to difficulty in identification and tracking
Solution Approach 1:
The patent segments sensitive data identification into two phases: (1) extraction of content from structured data entities into a linearized stream, and (2) classification of sensitive portions within that stream. This segmentation allows the system to handle distributed data across multiple cells and objects by processing it as a unified sequence while tracking original locations through offset information.
Solution Approach 2:
The patent introduces a content stream as an intermediary representation between the structured data entities and the sensitive data classification process. This content stream serves as a mediator that linearizes the distributed data while preserving location information through offsets, enabling effective DLP analysis without requiring direct manipulation of the complex structured formats.
2Device complexity
If traditional DLP methods are applied to structured documents, then implementation simplicity is maintained, but detection accuracy deteriorates because sensitive data may be split across multiple data entities
Solution Approach 1:
The patent transforms the detection problem from a two-dimensional structured format (rows and columns in spreadsheets, slides in presentations) to a one-dimensional linearized content stream. This dimensional transformation simplifies the classification process while maintaining the ability to accurately detect sensitive data through offset mapping back to the original structured locations.
3Reliability
If comprehensive sensitive data scanning is performed on all content, then data protection coverage is improved, but processing time increases due to the need to analyze all structured data entities
Solution Approach 1:
The patent extracts only the relevant content portions from the structured data entities into a linearized stream for classification analysis. By extracting content rather than processing entire structured documents, the system achieves comprehensive scanning of all potential sensitive data while significantly reducing processing time through more efficient data representation and analysis.
Data Source
AI summary
Systems, methods, and software for sensitive data handling frameworks for user applications are provided herein. An exemplary method includes receiving subsets of structured user content consolidated into associated flattened representations, the associated flattened representations having a mapping to the structured user content and accompanied by at least lengths and offset information relating to the mapping. The method includes individually parsing the subsets of structured user content to classify portions as comprising sensitive content corresponding to one or more predetermined data schemes and, for each of the portions, identifying an associated offset and length for the portion relating to the subsets of structured user content, and indicating at least the associated offset and length to the user application for marking of the sensitive content in a user interface to the user application.


